Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6d2ae104d7 | ||
|
|
b7fc543a83 | ||
|
|
b868033ddc | ||
|
|
8aad64cced | ||
|
|
319163ee7e | ||
|
|
87b16b7144 | ||
|
|
20ae6501f2 | ||
|
|
a16fe88114 | ||
|
|
c77598adc7 | ||
|
|
29f8fac013 | ||
|
|
6199e5a69e | ||
|
|
375b4b6bdb | ||
|
|
fdb3e42d6f | ||
|
|
51a6a16221 | ||
|
|
0798e24d24 | ||
|
|
3be2d08ac9 | ||
|
|
b7d5b23ae6 | ||
|
|
6eaa1c1a8b | ||
|
|
f341d22aa0 | ||
|
|
1a479ceb24 | ||
|
|
ff553d018a | ||
|
|
c807d48a13 | ||
|
|
9c383c1c8b | ||
|
|
462e995adc | ||
|
|
ea069b98da | ||
|
|
deb5419c47 |
+144
-63
@@ -5,11 +5,12 @@ Struct dumped to disk — separate from the kanta content database, path from
|
|||||||
`PAGERITE_ANALYTICS` (default: the database path with `.kantadb` replaced by
|
`PAGERITE_ANALYTICS` (default: the database path with `.kantadb` replaced by
|
||||||
`.analytics.json`, e.g. `pagerite.analytics.json`).
|
`.analytics.json`, e.g. `pagerite.analytics.json`).
|
||||||
|
|
||||||
- `pagerite/analytics.py` — data model (`Analytics`, `Visit`) and the `Store`
|
- `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`,
|
||||||
(in-memory data + session map, atomic JSON persistence).
|
`CrawlerHit`, `AbuseHit`) and the `Store` (in-memory data + session map,
|
||||||
|
atomic JSON persistence).
|
||||||
- `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`),
|
- `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`),
|
||||||
the `POST /_a` ping endpoint, and `GET /_api/analytics` (admin-gated like
|
the `POST /_a` ping endpoint, and `WebSocket /_api/ws/analytics`
|
||||||
every `/_api` endpoint).
|
(admin-gated like every `/_api` endpoint).
|
||||||
- `frontend/src/pagerite.js` — client navigation pings and the 📊 pen.
|
- `frontend/src/pagerite.js` — client navigation pings and the 📊 pen.
|
||||||
- `frontend/src/AnalyticsView.vue` — viewer component rendered inside the
|
- `frontend/src/AnalyticsView.vue` — viewer component rendered inside the
|
||||||
normal site layout on the `/_a` analytics page.
|
normal site layout on the `/_a` analytics page.
|
||||||
@@ -23,7 +24,14 @@ The client (`pagerite.js`) POSTs fire-and-forget pings to `/_a` with
|
|||||||
|
|
||||||
- **Initial page load**: `to` is the loaded path. This ping is what starts
|
- **Initial page load**: `to` is the loaded path. This ping is what starts
|
||||||
the visit and counts the entry page view — the document GET alone records
|
the visit and counts the entry page view — the document GET alone records
|
||||||
nothing, so bots and admin browsing never register. Reloads are not
|
nothing, so bots and admin browsing never register. JS-running crawlers
|
||||||
|
(Googlebot, GoogleOther, Applebot, ...) do ping, but their User-Agent
|
||||||
|
gives them away: pings whose UA matches `_is_bot_ua` (anything calling
|
||||||
|
itself a "bot", plus known exceptions such as GoogleOther) are ignored
|
||||||
|
server-side, and their document GETs land in the crawler list instead.
|
||||||
|
No source-IP verification is done: a spoofed bot UA merely lands in the
|
||||||
|
crawler stats, and scanners that probe telltale paths are caught by the
|
||||||
|
abuse rules regardless. Reloads are not
|
||||||
visits: the ping is skipped (PerformanceNavigationTiming `reload`), so a
|
visits: the ping is skipped (PerformanceNavigationTiming `reload`), so a
|
||||||
refresh neither counts a second view nor logs a self-transition. The GET
|
refresh neither counts a second view nor logs a self-transition. The GET
|
||||||
handler stashes a cross-origin https `Referer` (origin part only) and any
|
handler stashes a cross-origin https `Referer` (origin part only) and any
|
||||||
@@ -32,76 +40,138 @@ The client (`pagerite.js`) POSTs fire-and-forget pings to `/_a` with
|
|||||||
- **Internal fetch-navigations**: `to` is the target path, sent only after
|
- **Internal fetch-navigations**: `to` is the target path, sent only after
|
||||||
the swap actually happened (a failed swap falls back to a full load,
|
the swap actually happened (a failed swap falls back to a full load,
|
||||||
whose initial ping counts the view instead — no gap, no double count).
|
whose initial ping counts the view instead — no gap, no double count).
|
||||||
- **External links** (`https` only): `to` is the link's origin. This is the
|
- **External links** (`https` only): `to` is the link's full URL. This is the
|
||||||
exit-link record; the user may continue navigating afterwards (new tab,
|
exit-link record; the user may continue navigating afterwards (new tab,
|
||||||
back), so the exit origin is not necessarily the last trail entry.
|
back), so the exit URL is not necessarily the last trail entry. Outbound
|
||||||
|
links are stored by full URL so several links to the same domain remain
|
||||||
|
distinct.
|
||||||
- **Excluded**: back/forward (popstate) navigations, navigation involving
|
- **Excluded**: back/forward (popstate) navigations, navigation involving
|
||||||
the analytics page itself (`/_a`), and everything while the user is known to
|
the analytics page itself (`/_a`), and everything while the user has the
|
||||||
be an admin *and SSO is actually in use* — with no auth proxy (dev/test)
|
editor open (`body.editing`). Admin noise, not visits.
|
||||||
"admin" is everyone's state, so the gate is off and everything is recorded —
|
- **Admins**: when SSO is in use and the session is known to be an admin,
|
||||||
or has the editor open (`body.editing`). Admin noise, not visits.
|
the client still pings but adds `hide=1`. The server then records
|
||||||
|
nothing — and if the same client session already had a visit from before
|
||||||
|
logging in, that visit is removed from the JSON along with the counts
|
||||||
|
recorded when it was created (site visit, entry view, entry transition).
|
||||||
|
Views/transitions logged by later pings inside such a visit lack
|
||||||
|
per-event timestamps and are left as-is. With no auth proxy (dev/test)
|
||||||
|
"admin" is everyone's state, so `hide` stays 0 and everything is recorded.
|
||||||
- The server validates `to`: internal paths must be valid slug paths
|
- The server validates `to`: internal paths must be valid slug paths
|
||||||
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the
|
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the
|
||||||
https origin and accepted only when the client sent exactly that.
|
https origin and accepted only when the client sent exactly that.
|
||||||
- The initial ping also records the visitor's `User-Agent` and
|
- **Client records**: the visitor's IP (IPv4 or IPv6 /64 network), raw
|
||||||
`Accept-Language` headers. The first `Accept-Language` tag is stored as
|
`User-Agent` and extracted `Accept-Language` tag are hashed with blake3;
|
||||||
`lang` (e.g. `en-us`) and its region subtag, if present, is stored as
|
the first 6 bytes identify a shared `Client` record. The `Client` stores
|
||||||
an initial `country` (e.g. `US`).
|
the full IP, `User-Agent`, compact `ua_pretty`, `lang`, initial
|
||||||
- The visitor IP is stored. A reverse-DNS lookup is attempted for each new
|
`country` from the language-region subtag, and asynchronously-filled
|
||||||
visit and the result, when available, is cached in RAM and stored as
|
`country`/`city` from DB-IP geoip plus reverse-DNS `host`. Visits,
|
||||||
`host`; local/reserved/multicast addresses are skipped.
|
crawler hits and abuse hits all reference this record by its hash, so
|
||||||
- If a DB-IP MMDB file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the
|
client metadata is stored once instead of repeated per event.
|
||||||
repository root, it is loaded at startup and used to look up a more accurate
|
- The visitor IP is stored in the `Client`. A reverse-DNS lookup is
|
||||||
`country`. The MMDB lookup and the reverse-DNS lookup run in background
|
attempted for each new client and the result, when available, is stored as
|
||||||
tasks after the visit is stored, so the `/ _a` response is never delayed.
|
`host`; local/reserved/multicast addresses are skipped. If a DB-IP MMDB
|
||||||
The decompressed `dbip-*.mmdb` file is kept in the repository root and
|
file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the repository
|
||||||
ignored by git.
|
root, it is loaded at startup and used to look up `country`/`city`. These
|
||||||
|
lookups run in background tasks after the event is stored, so the `/_a`
|
||||||
|
response is never delayed. The decompressed `dbip-*.mmdb` file is kept in
|
||||||
|
the repository root and ignored by git. The CLI flag `--dbip`
|
||||||
|
(`uv run pagerite --dbip`) downloads the latest
|
||||||
|
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server starts,
|
||||||
|
skipping the download when the local database is already current and
|
||||||
|
removing older versions after an update; without the flag only an existing
|
||||||
|
file is used.
|
||||||
- **Crawler hits**: every document GET is queued in RAM as a pending crawler
|
- **Crawler hits**: every document GET is queued in RAM as a pending crawler
|
||||||
hit. If a ping from the same (IP, User-Agent) pair arrives within 10
|
hit — except idle-time link preloads from pagerite.js, which carry an
|
||||||
seconds the hit is discarded; otherwise it is written to `crawlers`.
|
`x-pagerite-preload` header and are not tracked at all (the ping sent when
|
||||||
Crawlers do not count as visits or views.
|
the user actually navigates to a preloaded page does the counting; forging
|
||||||
|
the header only hides a GET from the crawler stats, the path-based abuse
|
||||||
|
classification is unaffected). If a ping
|
||||||
|
from the same client arrives within 10 seconds the hit is discarded;
|
||||||
|
otherwise it is written to `crawlers`. Crawlers do not count as
|
||||||
|
visits or views. The `Accept-Language` header is stored on the shared
|
||||||
|
`Client` immediately; reverse-DNS host names and DB-IP geoip
|
||||||
|
country/city are filled in asynchronously, just like for real visits. In
|
||||||
|
the analytics viewer, crawler hits are grouped by client hash and shown as
|
||||||
|
a trail of internal pages that crawler visited; the crawler table lists
|
||||||
|
the most active crawlers first rather than the most recent hits.
|
||||||
|
- **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment
|
||||||
|
starting with a dot (`/.env`, `/.git/config`) or ending in `.php` —
|
||||||
|
classifies the source IP as abuse immediately, and ten plain 404s from one
|
||||||
|
IP do too. Classification reclassifies history: all earlier crawler hits
|
||||||
|
from that IP (persisted and pending) move to the `abuse` list, so a
|
||||||
|
random-UA scanner no longer pollutes the crawler stats of the legitimate
|
||||||
|
bot it impersonates. Once classified, every document GET and 404 from the
|
||||||
|
IP is recorded as an abuse hit with the full request path (query string
|
||||||
|
included), and its pings are ignored. The classified IP set (`abuse_ips`)
|
||||||
|
is persisted in the JSON file; the plain-404 counters are RAM-only. In the
|
||||||
|
viewer, abuse hits are grouped by IP (never by client/UA — scanners
|
||||||
|
randomize theirs) in a separate "Abuse" table. Identical paths are
|
||||||
|
collapsed into one entry with their hit count; flagged paths that
|
||||||
|
triggered classification are lifted to the top, followed by other 404s and
|
||||||
|
then document GETs from the abuser. Raw User-Agent strings are shown one
|
||||||
|
per line with their occurrence counts, and the full lists are click-to-copy.
|
||||||
|
|
||||||
## Visits and sessions
|
## Visits and sessions
|
||||||
|
|
||||||
There are no cookies. A visit is tied together by the (IP, User-Agent) pair
|
There are no cookies. A visit is tied together by a client hash — the first
|
||||||
(IP from the first `X-Forwarded-For` hop — we sit behind a proxy — else the
|
6 bytes of a blake3 digest over the prettified IP (IPv4 unchanged, IPv6
|
||||||
direct peer): the first ping from a pair starts a new visit, subsequent
|
/64 network), the raw `User-Agent` string and the extracted
|
||||||
pings extend it. Pings arriving with no known session (server restart)
|
`Accept-Language` tag. The first ping from a client hash starts a new
|
||||||
start a fresh visit from the first ping — treated as missing data rather
|
visit; subsequent pings extend it. Pings arriving with no known session
|
||||||
than dropped. The (IP, UA) → visit map and the IP → entry-referer/UTM
|
(server restart) start a fresh visit from the first ping — treated as
|
||||||
tables are in-memory only, but the IP and any resolvable reverse-DNS host
|
missing data rather than dropped. The client-hash → visit map and the IP →
|
||||||
name are stored on the `Visit` record itself.
|
entry-referer/UTM tables are in-memory only; client metadata is stored in
|
||||||
|
`Analytics.clients` keyed by the client hash.
|
||||||
|
|
||||||
|
Each `Client` record:
|
||||||
|
|
||||||
|
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
|
||||||
|
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
|
||||||
|
- `lang` — first `Accept-Language` tag, lowercased (e.g. `"en-us"`),
|
||||||
|
- `country` — two-letter country code. Initially derived from the
|
||||||
|
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
|
||||||
|
when a database is available,
|
||||||
|
- `city` — city name from the DB-IP MMDB lookup, when available,
|
||||||
|
- `ua` — raw `User-Agent` string,
|
||||||
|
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
|
||||||
|
parsable, otherwise the raw string.
|
||||||
|
|
||||||
Each `Visit` record:
|
Each `Visit` record:
|
||||||
|
|
||||||
- `start` — timestamp of the first event,
|
- `start` — timestamp of the first event,
|
||||||
- `entry` — first page (path) seen,
|
- `entry` — first page (path) seen,
|
||||||
- `referer` — external https origin of the initial load, `""` for direct,
|
- `referer` — external https origin of the initial load, `""` for direct,
|
||||||
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
|
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
|
||||||
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
|
|
||||||
- `trail` — everything seen afterwards in first-seen order: page paths and
|
- `trail` — everything seen afterwards in first-seen order: page paths and
|
||||||
external exit origins. Re-visiting an already seen page (incl. the entry)
|
external exit URLs. Re-visiting an already seen page (incl. the entry)
|
||||||
does not append.
|
does not append.
|
||||||
- `lang` — first `Accept-Language` tag, lowercased (e.g. `en-us`),
|
|
||||||
- `country` — two-letter country code. Initially derived from the
|
|
||||||
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
|
|
||||||
when a database is available,
|
|
||||||
- `ua` — raw `User-Agent` string from the initial ping,
|
|
||||||
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
|
|
||||||
parsable, otherwise the raw string,
|
|
||||||
- `utm` — `utm_*` query parameters from the landing URL, as a dict.
|
- `utm` — `utm_*` query parameters from the landing URL, as a dict.
|
||||||
|
- `read` — active reading time per path (seconds), keyed by path.
|
||||||
|
|
||||||
Each `CrawlerHit` record:
|
Each `CrawlerHit` record:
|
||||||
|
|
||||||
- `start` — timestamp of the document GET,
|
- `start` — timestamp of the document GET,
|
||||||
- `entry` — page path requested,
|
- `entry` — page path requested,
|
||||||
- `ip` — IP address,
|
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
|
||||||
- `ua` — raw `User-Agent` header,
|
|
||||||
- `ua_pretty` — compact display form of the UA when parsable,
|
|
||||||
- `referer` — external https origin of the request, `""` for direct/none,
|
- `referer` — external https origin of the request, `""` for direct/none,
|
||||||
- `query` — raw query string of the request.
|
- `query` — raw query string of the request.
|
||||||
|
|
||||||
Crawler hits are grouped by User-Agent in the analytics viewer.
|
Each `AbuseHit` record:
|
||||||
|
|
||||||
|
- `start` — timestamp of the request,
|
||||||
|
- `path` — full request path including the query string (e.g. `/.env?x=1`),
|
||||||
|
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
|
||||||
|
- `flag` — true for the path that triggered abuse classification (telltale
|
||||||
|
path or the 404 that crossed the threshold),
|
||||||
|
- `is_404` — true for 404 responses, false for document GETs from the
|
||||||
|
abuser.
|
||||||
|
|
||||||
|
Crawler hits are grouped by client hash in the analytics viewer; abuse hits
|
||||||
|
are grouped by IP alone (resolved from the referenced `Client`). In the
|
||||||
|
Abuse table identical paths are collapsed with their counts; flagged paths
|
||||||
|
that triggered classification are lifted to the top, followed by other 404s
|
||||||
|
and then document GETs from the abuser. Within each category paths are
|
||||||
|
sorted by count descending, then by their earliest hit.
|
||||||
|
|
||||||
## Aggregates
|
## Aggregates
|
||||||
|
|
||||||
@@ -131,15 +201,17 @@ The 📊 pen in the banner corner (admins only, injected by pagerite.js next to
|
|||||||
the edit pens) links to `/_a`, the analytics page. It is a normal site page:
|
the edit pens) links to `/_a`, the analytics page. It is a normal site page:
|
||||||
the standard banner, navigation and footer stay in place, and the analytics
|
the standard banner, navigation and footer stay in place, and the analytics
|
||||||
content is rendered inside `#main`. The page itself is public, but the data
|
content is rendered inside `#main`. The page itself is public, but the data
|
||||||
still comes from `GET /_api/analytics`, which remains admin-gated like the
|
stream comes from `WebSocket /_api/ws/analytics`, which remains admin-gated
|
||||||
rest of the management API; visitors without access see the viewer with a
|
like the rest of the management API; visitors without access see the viewer
|
||||||
"could not be loaded" message.
|
with a "could not be loaded" message.
|
||||||
|
|
||||||
Because it is a real page, fetch-navigation handles it like any other internal
|
Because it is a real page, fetch-navigation handles it like any other internal
|
||||||
link: clicking the 📊 pen (or any link to `/_a`) fetches the server-rendered
|
link: clicking the 📊 pen (or any link to `/_a`) fetches the server-rendered
|
||||||
HTML, swaps the dynamic regions and mounts the Vue analytics app in place. The
|
HTML, swaps the dynamic regions and mounts the Vue analytics app in place. The
|
||||||
range selector updates the URL query string (`?range=week` etc.) so links to
|
range selector updates the URL hash (`#week` etc.) so links to a specific
|
||||||
a specific range can be shared.
|
range can be shared. When the URL has no hash, the client derives the
|
||||||
|
default from the first analytics snapshot: `day` if the recorded history
|
||||||
|
spans less than 24 hours, otherwise `week`.
|
||||||
|
|
||||||
`AnalyticsView.vue` is no longer a full-screen overlay; the `body.analytics-open`
|
`AnalyticsView.vue` is no longer a full-screen overlay; the `body.analytics-open`
|
||||||
page-chrome hiding and `#/analytics/<range>` hash routing have been removed.
|
page-chrome hiding and `#/analytics/<range>` hash routing have been removed.
|
||||||
@@ -156,16 +228,19 @@ the smoothing time scale follows the unit: the month+ sigmas are 24× the
|
|||||||
hourly ones. The y max is derived from the smoothed curves so single-bucket
|
hourly ones. The y max is derived from the smoothed curves so single-bucket
|
||||||
spikes don't blow up the scale, and raw spikes are clamped into the plot.
|
spikes don't blow up the scale, and raw spikes are clamped into the plot.
|
||||||
Axes always start at 0 and end at a multiple of a 1-2-5 major step (max 5
|
Axes always start at 0 and end at a multiple of a 1-2-5 major step (max 5
|
||||||
labeled intervals, minor lines at fifths when integral; the floor is 1/h).
|
labeled intervals, minor lines at fifths when integral; the minimum y-axis
|
||||||
|
range is 10 so tiny values such as a single visit are not stretched to a
|
||||||
|
fractional scale).
|
||||||
The week range is aligned to Monday 00:00 UTC and overlays up to 8 previous
|
The week range is aligned to Monday 00:00 UTC and overlays up to 8 previous
|
||||||
weeks in the same accent color at decreasing opacity (the current week is
|
weeks in the same accent color at decreasing opacity (the current week is
|
||||||
truncated at the current bucket, never drawing fake zeroes for the future);
|
truncated at the current bucket, never drawing fake zeroes for the future);
|
||||||
its x labels are weekday names centered at midday UTC, without vertical grid
|
its x labels are weekday names centered at midday UTC, without vertical grid
|
||||||
lines (day boundaries would be misleading in the viewer's timezone). The
|
lines (day boundaries would be misleading in the viewer's timezone). The
|
||||||
month view labels days the same lineless way — day numbers at noon UTC,
|
month view labels days the same lineless way — day numbers at noon UTC,
|
||||||
with the month name substituted for the 1st. Year and all are rolling
|
with the month name substituted for the 1st. Year is a rolling 365-day window ending at now, re-bucketed to daily points,
|
||||||
windows ending at now, re-bucketed to daily points, with boundary lines at
|
with boundary lines at months/years. All uses the full data reach, but keeps
|
||||||
months/years. Below the charts: a radial **transition map** (all pages from
|
at least the past 30 days so the chart never collapses to a tiny sliver when
|
||||||
|
the site is young. Below the charts: a radial **transition map** (all pages from
|
||||||
`/_api/pages` — front page at the center, each slug level on its own ring,
|
`/_api/pages` — front page at the center, each slug level on its own ring,
|
||||||
siblings clockwise in navigation order from the top, radial gap equal to
|
siblings clockwise in navigation order from the top, radial gap equal to
|
||||||
the arc spacing — opposite transition directions joined into organic
|
the arc spacing — opposite transition directions joined into organic
|
||||||
@@ -175,8 +250,14 @@ carrying less than 1% of the total traffic
|
|||||||
pruned; beads are simulated one by one in JS (requestAnimationFrame) and
|
pruned; beads are simulated one by one in JS (requestAnimationFrame) and
|
||||||
flow along each edge, emitted at a rate linearly proportional
|
flow along each edge, emitted at a rate linearly proportional
|
||||||
to the directional count with no in-flight limit, opposing directions
|
to the directional count with no in-flight limit, opposing directions
|
||||||
offset onto parallel lanes. External referers show as a node row above the
|
offset onto parallel lanes. External sources show as a node row above the
|
||||||
map, external exits as small nodes fanned outwards from their source
|
map: each visit is attributed to `utm_campaign`, then `utm_source`, then the
|
||||||
page), per-page view
|
referer origin, then any other `utm_*` tag, so UTM-tagged visits are grouped
|
||||||
counts, the top transitions and the 50 most recent visit trails. Data comes from `GET /_api/analytics`, which
|
under their campaign/source value rather than the referer domain. A UTM
|
||||||
returns the raw JSON file contents.
|
source node only links to its referer when every visit carrying that tag
|
||||||
|
came from the same origin. External exits are small nodes fanned outwards
|
||||||
|
from their source page), per-page view
|
||||||
|
counts, the top transitions and the 50 most recent visit trails. Data is
|
||||||
|
streamed live over `WebSocket /_api/ws/analytics`, which pushes the latest
|
||||||
|
JSON snapshot on connect and again whenever the analytics file is updated
|
||||||
|
(with a small server-side debounce to avoid flooding under high traffic).
|
||||||
|
|||||||
+3
-1
@@ -8,6 +8,8 @@ The FastAPI app. FastAPI's built-in API docs are disabled (`docs_url`/`redoc_url
|
|||||||
|
|
||||||
The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours.
|
The build mirrors the URL space — hashed immutable assets under `/_assets/`, `favicon.ico` at the site root — and an `index.html` in the build would become a `/` route, so leave it out of the build to keep `/` ours.
|
||||||
|
|
||||||
|
Generated HTML pages (content pages, category/404 placeholders, `/_a`) go through `_html_response`: zstd-compressed per request at level 9 when the client sends `accept-encoding: zstd` (no gzip fallback; static assets are pre-compressed by the `Frontend`), with `vary: accept-encoding` set and the ETag kept identical across encodings so `if-none-match` revalidation still works. In production the rendered bodies are cached in an LRU keyed by everything the output depends on — page kind, path, the site origin (social meta), encoding, and `data.version`, which bumps on every content/settings change and so transparently invalidates the whole cache. The cache is bypassed in dev, where theme/design CSS is re-read from disk per request. Content pages carry an ETag built from the node's modified timestamp and `data.version`; `/_a` instead gets a blake3 hash of the rendered body (it has no Node), with matching `if-none-match` revalidations answered by a 304.
|
||||||
|
|
||||||
## `data.py`
|
## `data.py`
|
||||||
|
|
||||||
msgspec Structs for the kanta database. See `docs/content-model.md` for the full data model.
|
msgspec Structs for the kanta database. See `docs/content-model.md` for the full data model.
|
||||||
@@ -20,7 +22,7 @@ markdown-it-py renderer (html passthrough + attrs, footnote, deflist, tasklists,
|
|||||||
|
|
||||||
The shared page layout as an html5tagger `Template` with placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav rendering straight from the `Data.menu` tree (siblings sorted by `Node.order`; nav links to content-less labels point at their first child via `first_leaf`, the first published descendant with content), and page/404 rendering.
|
The shared page layout as an html5tagger `Template` with placeholders (`Title`, `Brand`, `Banner`, `Nav`, `Sidebar`, `Main`), nav rendering straight from the `Data.menu` tree (siblings sorted by `Node.order`; nav links to content-less labels point at their first child via `first_leaf`, the first published descendant with content), and page/404 rendering.
|
||||||
|
|
||||||
Content pages get SEO/social meta (description, canonical link, Open Graph + twitter card) from heuristics over the rendered article: the description is the first paragraph's text, the share image prefers a `{.hero}`-classed image, then the first raster `<img>`, then the first SVG; the first `<video>` yields `og:video`; URLs are made absolute with the request base URL; `article:published/modified_time` come from `Node.created`/`modified`. If the markdown contains its own h1, the page title is NOT rendered as an additional h1 (it still supplies `<title>` and nav labels).
|
Content pages get SEO/social meta (description, canonical link, Open Graph + twitter card) from heuristics over the rendered article: the description is the first paragraph's text, the share image prefers a `{.hero}`-classed image, then the first raster `<img>`, then the first SVG; the first `<video>` yields `og:video`; URLs are made absolute with the site origin (`Data.site_url` — learned from admin browsers reporting their `location.origin` via `POST /_api/site-url`, correct even behind reverse proxies; until learned, the request's own base URL is the fallback); `article:published/modified_time` come from `Node.created`/`modified`. If the markdown contains its own h1, the page title is NOT rendered as an additional h1 (it still supplies `<title>` and nav labels).
|
||||||
|
|
||||||
The navbar holds top-level items only; the current section's subitems go to a left `#sidebar` as a nested list (the section's direct children plain, deeper levels indented with article-list-style markers), which is rendered when the section offers at least two published items, or exactly one while viewing anything other than that only page — the section index, a 404, a grandchild (so those pages can reach the child), and also on that only page itself when it has published children of its own; no aside element at all on the front page, leaf pages and the sole childless page of a one-page section. Also, category labels are nodes without content — None *or* empty markdown — and their nav links point at their first child page. Dynamic regions have stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps (`#sidebar` may be absent on either side of a swap).
|
The navbar holds top-level items only; the current section's subitems go to a left `#sidebar` as a nested list (the section's direct children plain, deeper levels indented with article-list-style markers), which is rendered when the section offers at least two published items, or exactly one while viewing anything other than that only page — the section index, a 404, a grandchild (so those pages can reach the child), and also on that only page itself when it has published children of its own; no aside element at all on the front page, leaf pages and the sole childless page of a one-page section. Also, category labels are nodes without content — None *or* empty markdown — and their nav links point at their first child page. Dynamic regions have stable ids (`#page-banner`, `#nav`, `#sidebar`, `#main`) for fetch-navigation swaps (`#sidebar` may be absent on either side of a swap).
|
||||||
|
|
||||||
|
|||||||
@@ -20,7 +20,7 @@ Siblings order by the fractional `Node.order` key: a moved item gets a fresh key
|
|||||||
|
|
||||||
`Node.banner` is a raw trusted HTML snippet for the header banner (img, styled div, canvas+script...); empty inherits from the node's ancestors (front page last). It is rendered AFTER the banner design's artwork, so author code (e.g. a `<style>` override) always wins over the design's own styles.
|
`Node.banner` is a raw trusted HTML snippet for the header banner (img, styled div, canvas+script...); empty inherits from the node's ancestors (front page last). It is rendered AFTER the banner design's artwork, so author code (e.g. a `<style>` override) always wins over the design's own styles.
|
||||||
|
|
||||||
`Node.banner_design` picks a banner design: a theme folder name whose `banner.css` styles it and whose `banner.html` (arbitrary markup: canvas + style + script) or `banner.svg` supplies the inline artwork (wrapped in `div[data-design]`); "" = explicitly no design, None = inherit (nearest ancestor, front page last, then the active theme's own design if it ships banner.css/banner.svg/banner.html). The design's banner.css is linked in `<head>` (id `pagerite-banner`) between the theme and the custom CSS.
|
`Node.banner_design` picks a banner design: a theme folder name whose `banner.css` styles it and whose `banner.html` (arbitrary markup: canvas + style + script) or `banner.svg` supplies the inline artwork (wrapped in `div[data-design]`); "" = explicitly no design, None = inherit (nearest ancestor, front page last, then the active theme's own design if it ships banner.css/banner.svg/banner.html). The design's banner.css lives in `<head>` (id `pagerite-banner`) between the theme and the custom CSS — a `<link>` in dev, an inline `<style>` in production.
|
||||||
|
|
||||||
## Site settings
|
## Site settings
|
||||||
|
|
||||||
|
|||||||
+1
-1
@@ -25,6 +25,6 @@ Dropping ON the lower part of a row moves the page under that row (the child lis
|
|||||||
|
|
||||||
The shell is dynamic-imported onto the content page by pagerite.js when an edit pen is clicked (the pens are injected by pagerite.js after the session validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`). In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`), in prod from the hashed build assets resolved via `frontend-build/.vite/manifest.json`.
|
The shell is dynamic-imported onto the content page by pagerite.js when an edit pen is clicked (the pens are injected by pagerite.js after the session validates; they carry `data-editor-src`/`data-editor-css`/`data-editor-mode`). In dev, modules load from the Vite dev server (`PAGERITE_VITE_URL`), in prod from the hashed build assets resolved via `frontend-build/.vite/manifest.json`.
|
||||||
|
|
||||||
`vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_/assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js` and `src/pagerite.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme and banner-design CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-user` elements to restore the canonical order (base < theme < design < custom CSS). Theme switches in the site editor simply swap the `#pagerite-theme` link href, identically in dev and prod.
|
`vite.config.js` sets `appType: 'mpa'` (no SPA fallback) and builds with `manifest: true`, `assetsDir: '_/assets'` (so the build mirrors the URL space; `frontend/public/favicon.ico` lands at the build root and is served at `/favicon.ico`). JS inputs are `src/main.js` and `src/pagerite.js`, plus `src/assets/pagerite.css` as a separate stylesheet entry; theme and banner-design CSS are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. There is no `index.html` source (it would shadow `/` and turn missing dev paths into an empty Vue shell). All outputs are ES modules. The build sets `preserveEntrySignatures: 'exports-only'` because main.js is consumed via dynamic `import()` for its `openEditor`/`closeEditor` exports — Vite app builds otherwise strip unused entry exports, leaving dead edit pens. In dev the backend links theme/banner-design stylesheets like in prod (`/_themes/...`); only the base CSS is Vite-injected from JS, and pagerite.js then re-appends the `#pagerite-theme`/`#pagerite-banner`/`#pagerite-user` elements to restore the canonical order (base < theme < design < custom CSS). In production all page assets are inlined instead (styles as `<style id="pagerite-…">` in `<head>`, scripts at the end of the body). Theme switches in the site editor swap the `#pagerite-theme` element in place — the link href in dev, the inline style's text (fetched from `/_themes/...`) in prod.
|
||||||
|
|
||||||
`vite-plugin-fastapi.js` has an auto-upgrade marker — edit `vite.config.js`, not the plugin.
|
`vite-plugin-fastapi.js` has an auto-upgrade marker — edit `vite.config.js`, not the plugin.
|
||||||
|
|||||||
@@ -8,9 +8,11 @@ Vue editor app entry, mounts the tabbed `EditorShell`. See `docs/editing.md` for
|
|||||||
|
|
||||||
## `pagerite.js`
|
## `pagerite.js`
|
||||||
|
|
||||||
Public page entry; runs fetch-navigation (backed by an in-memory page cache: every visible internal link — and the current page — is fetched once at load, clicks are then served from JS with no fetch, and the editors' `loadPlain` keeps the cache current via a `pagerite:page-fetched` event; articles are `cache-control: no-cache` on the wire), scroll-reveal, OverlayScrollbars on `document.body` (floating, auto-hiding scrollbars that never reserve layout space or shift the page when appearing; native scroll APIs like `window.scrollTo` keep working; themed via the `--os-*` variables in pagerite.css), brand shrink-to-fit (the themed size is the maximum; JS reduces the font-size so a long brand or narrow viewport still fits one line), code copy buttons, and the auth check.
|
Public page entry; runs fetch-navigation (backed by an in-memory page cache: every visible internal link is fetched once at load and clicks are then served from JS with no fetch — the current page itself is not refetched, it enters the cache when navigated to — and the editors' `loadPlain` keeps the cache current via a `pagerite:page-fetched` event; articles are `cache-control: no-cache` on the wire), scroll-reveal, OverlayScrollbars on `document.body` (floating, auto-hiding scrollbars that never reserve layout space or shift the page when appearing; native scroll APIs like `window.scrollTo` keep working; themed via the `--os-*` variables in pagerite.css), brand shrink-to-fit (the themed size is the maximum; JS reduces the font-size so a long brand or narrow viewport still fits one line), nav condense-to-fit (the top nav stays on one row: link gaps shrink first, then the side padding, then the font size; `flex-wrap: wrap` remains the no-JS fallback), code copy buttons, and the auth check.
|
||||||
|
|
||||||
It first probes `GET /auth/api/settings` to detect whether Paskia SSO is available, then `GET /_api/settings` to learn the current session's admin status. The same reverse proxy that gates `/_api` returns 401 for anonymous users, 403 for users without the admin permission, and 200 for admins. When Paskia is detected, a login link (anonymous) or profile link (logged in) is shown in the banner corner; both are plain `<a href="/auth/">` links (Paskia does not support being iframed, so we navigate normally), and a `pageshow` handler re-probes auth when history navigation restores a cached page. Admins also get the page/banner edit pens and a site-settings pen (asset URLs from the `pagerite:editor-src`/`-css` meta tags). If no Paskia SSO is detected (dev/no proxy), editing is left open. Pages themselves render identically for everyone; the real gate is the auth proxy in front of all of `/_api`. The backend links the stylesheets in a fixed order — base (Vite build), theme, banner design, custom CSS last — each with a stable id so the site editor can swap them in place.
|
It first probes `GET /auth/api/settings` to detect whether Paskia SSO is available, then `GET /_api/settings` to learn the current session's admin status. The same reverse proxy that gates `/_api` returns 401 for anonymous users, 403 for users without the admin permission, and 200 for admins. When Paskia is detected, a login link (anonymous) or profile link (logged in) is shown in the banner corner; both are plain `<a href="/auth/">` links (Paskia does not support being iframed, so we navigate normally), and a `pageshow` handler re-probes auth when history navigation restores a cached page. Admins also get the page/banner edit pens and a site-settings pen, plus a `modulepreload` warm-up of the editor bundle (the hashed asset is immutable, so it costs nothing). If no Paskia SSO is detected (dev/no proxy), editing is left open. Pages themselves render identically for everyone; the real gate is the auth proxy in front of all of `/_api`.
|
||||||
|
|
||||||
|
Asset wiring differs by mode. In dev the backend links the Vite dev-server URLs (`pagerite:editor-src`/`-css`/`pagerite:analytics-src` meta tags, `<link>` stylesheets) and Vite injects the entry CSS from JS for hot reloads. In production there are no pagerite meta tags: all page assets are inlined into the document — stylesheets as `<style>` elements in `<head>` (fixed order: base, theme, banner design, entry sheets, custom CSS last), module scripts as inline `<script>`s at the end of the body (relative chunk imports are rewritten to absolute `/_assets/` paths) — and the on-demand bundles' URLs ride in a `<script type="application/json" id="pagerite-assets">` config. The editor bundle always stays external, imported on demand when a pen is opened. Every stylesheet element carries a stable id so fetch-navigation and the site editor can sync `<head>` positionally across swaps (the analytics sheet exists on `/_a` only and is added/removed as you navigate). The analytics entry is inlined into the `/_a` page itself; pagerite.js re-creates that script element after fetch-navigating there (inline scripts don't execute on a DOM swap) and calls the module's exposed unmount before swapping away.
|
||||||
|
|
||||||
## `assets/`
|
## `assets/`
|
||||||
|
|
||||||
@@ -18,7 +20,7 @@ Shared styles and data files built by Vite and served hashed under `/_assets/`:
|
|||||||
|
|
||||||
The `::view-transition*` block at the end of `pagerite.css` (from termotohtori.fi) is fragile — do not tweak. Themes and banner designs are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. See `docs/themes-and-assets.md` for details.
|
The `::view-transition*` block at the end of `pagerite.css` (from termotohtori.fi) is fragile — do not tweak. Themes and banner designs are NOT built — they live in `pagerite/themes/{name}/` and are served by the backend. See `docs/themes-and-assets.md` for details.
|
||||||
|
|
||||||
Vite builds ES-module `.js` outputs; the backend renders `<script type="module">` for them (module scripts defer by default).
|
Vite builds ES-module `.js` outputs; in dev the backend links them as `<script type="module">` (module scripts defer by default), in production it inlines them at the end of the body.
|
||||||
|
|
||||||
## Database file
|
## Database file
|
||||||
|
|
||||||
|
|||||||
@@ -34,4 +34,4 @@ The banner artwork has scroll parallax: pagerite.js sets the `--pry` scroll para
|
|||||||
|
|
||||||
## Stylesheet order
|
## Stylesheet order
|
||||||
|
|
||||||
The backend links the stylesheets in a fixed order — base (Vite build), theme, banner design, custom CSS last — each with a stable id so the site editor can swap them in place. The base stylesheet's `--font-brand` defaults to `var(--font-heading)`.
|
The backend emits the stylesheets in a fixed order — base (Vite build), theme, banner design, entry sheets, custom CSS last — each with a stable id so fetch-navigation and the site editor can sync them in place. In dev they are `<link>`s (the base is Vite-injected from JS instead); in production they are inlined as `<style>` elements. The base stylesheet's `--font-brand` defaults to `var(--font-heading)`.
|
||||||
|
|||||||
+229
-121
@@ -1,76 +1,112 @@
|
|||||||
<script setup>
|
<script setup>
|
||||||
// Analytics viewer rendered as a normal page inside #main. Fetches the raw
|
// Analytics viewer rendered as a normal page inside #main. Receives live
|
||||||
// collected data from /_api/analytics (admin-gated by the auth proxy) and
|
// analytics data over /_api/ws/analytics (admin-gated by the auth proxy) and
|
||||||
// renders totals, smoothed visit/views curves, a transition map, and recent
|
// renders totals, smoothed visit/views curves, a transition map, and recent
|
||||||
// visit/crawler tables. Read-only.
|
// visit/crawler tables. Read-only.
|
||||||
// See docs/analytics.md for the data format.
|
// See docs/analytics.md for the data format.
|
||||||
import { computed, onMounted, ref, watch } from 'vue'
|
import { computed, onMounted, onUnmounted, ref, watch } from 'vue'
|
||||||
import { RANGES } from './analytics/time.js'
|
import { RANGES } from './analytics/time.js'
|
||||||
import {
|
import {
|
||||||
|
calcReadStats,
|
||||||
calcTotalViews,
|
calcTotalViews,
|
||||||
copyIp,
|
copyIp,
|
||||||
countCrawlerUas,
|
copyList,
|
||||||
formatCounts,
|
formatCount,
|
||||||
|
formatAbuseRows,
|
||||||
formatCrawlerRows,
|
formatCrawlerRows,
|
||||||
formatVisitRows,
|
formatVisitRows,
|
||||||
} from './analytics/format.js'
|
} from './analytics/format.js'
|
||||||
import * as flagSvgs from 'country-flag-icons/string/3x2'
|
import TrailLink from './TrailLink.vue'
|
||||||
|
import VisitorCell from './VisitorCell.vue'
|
||||||
import TransitionGraph from './TransitionGraph.vue'
|
import TransitionGraph from './TransitionGraph.vue'
|
||||||
import VisitorCharts from './VisitorCharts.vue'
|
import VisitorCharts from './VisitorCharts.vue'
|
||||||
|
|
||||||
const props = defineProps({
|
const ABUSE_MAX_LINES = 5
|
||||||
initialRange: { type: String, default: 'week' },
|
|
||||||
})
|
|
||||||
|
|
||||||
const data = ref(null)
|
const data = ref(null)
|
||||||
const pageTree = ref(null)
|
const pageTree = ref(null)
|
||||||
const error = ref('')
|
const error = ref('')
|
||||||
|
const now = ref(Date.now())
|
||||||
|
let ws = null
|
||||||
|
let reconnectTimeout = null
|
||||||
|
let timeInterval = null
|
||||||
|
|
||||||
onMounted(async () => {
|
// The initial range comes from the URL hash (shareable links); without one,
|
||||||
try {
|
// it is derived from the first analytics snapshot: day when the recorded
|
||||||
const res = await fetch('/_api/analytics')
|
// history is shorter than 24 h, week otherwise.
|
||||||
if (!res.ok) throw new Error(res.statusText)
|
const hashRange = location.hash.slice(1)
|
||||||
data.value = await res.json()
|
const range = ref(RANGES[hashRange] ? hashRange : 'week')
|
||||||
} catch {
|
let rangePinned = Boolean(RANGES[hashRange])
|
||||||
|
|
||||||
|
function connectAnalytics() {
|
||||||
|
if (ws) return
|
||||||
|
const proto = location.protocol === 'https:' ? 'wss:' : 'ws:'
|
||||||
|
ws = new WebSocket(`${proto}//${location.host}/_api/ws/analytics`)
|
||||||
|
ws.onopen = () => { error.value = '' }
|
||||||
|
ws.onmessage = (event) => {
|
||||||
|
try {
|
||||||
|
data.value = JSON.parse(event.data)
|
||||||
|
if (!rangePinned) {
|
||||||
|
rangePinned = true
|
||||||
|
const starts = (data.value?.visits || [])
|
||||||
|
.map((v) => Date.parse(v.start))
|
||||||
|
.filter((t) => !Number.isNaN(t))
|
||||||
|
if (starts.length && Date.now() - Math.min(...starts) < 24 * 3600 * 1000) {
|
||||||
|
range.value = 'day'
|
||||||
|
}
|
||||||
|
}
|
||||||
|
} catch {
|
||||||
|
error.value = 'analytics data could not be loaded'
|
||||||
|
}
|
||||||
|
}
|
||||||
|
ws.onerror = () => {
|
||||||
error.value = 'analytics data could not be loaded'
|
error.value = 'analytics data could not be loaded'
|
||||||
}
|
}
|
||||||
|
ws.onclose = () => {
|
||||||
|
ws = null
|
||||||
|
reconnectTimeout = setTimeout(connectAnalytics, 2000)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
onMounted(async () => {
|
||||||
|
connectAnalytics()
|
||||||
|
now.value = Date.now()
|
||||||
|
timeInterval = setInterval(() => { now.value = Date.now() }, 1000)
|
||||||
// The site tree for the transition map (all pages in menu order). Not
|
// The site tree for the transition map (all pages in menu order). Not
|
||||||
// fatal: without it the map falls back to transition endpoints only.
|
// fatal: without it the map just narrows to pages seen in transitions.
|
||||||
try {
|
try {
|
||||||
const res = await fetch('/_api/pages')
|
const res = await fetch('/_api/pages')
|
||||||
if (res.ok) pageTree.value = await res.json()
|
if (res.ok) pageTree.value = await res.json()
|
||||||
} catch { /* map just narrows to pages seen in transitions */ }
|
} catch { /* map just narrows to pages seen in transitions */ }
|
||||||
})
|
})
|
||||||
|
|
||||||
|
onUnmounted(() => {
|
||||||
|
if (reconnectTimeout) clearTimeout(reconnectTimeout)
|
||||||
|
if (timeInterval) clearInterval(timeInterval)
|
||||||
|
if (ws) {
|
||||||
|
ws.onclose = null
|
||||||
|
ws.close()
|
||||||
|
ws = null
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
const visits = computed(() => data.value?.visits || [])
|
const visits = computed(() => data.value?.visits || [])
|
||||||
const totalViews = computed(() => calcTotalViews(data.value?.views))
|
const totalViews = computed(() => calcTotalViews(data.value?.views))
|
||||||
|
const readStats = computed(() => calcReadStats(visits.value))
|
||||||
const range = ref(RANGES[props.initialRange] ? props.initialRange : 'week')
|
|
||||||
|
|
||||||
// Keep the URL shareable when the range changes.
|
// Keep the URL shareable when the range changes.
|
||||||
watch(range, (r) => {
|
watch(range, (r) => {
|
||||||
const url = new URL(location.href)
|
const url = new URL(location.href)
|
||||||
url.searchParams.set('range', r)
|
url.hash = r
|
||||||
history.replaceState(null, '', url)
|
history.replaceState(null, '', url)
|
||||||
})
|
})
|
||||||
|
|
||||||
const visitRows = computed(() => formatVisitRows(visits.value, pageTree.value))
|
const clients = computed(() => data.value?.clients || {})
|
||||||
|
const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value))
|
||||||
const crawlers = computed(() => data.value?.crawlers || [])
|
const crawlers = computed(() => data.value?.crawlers || [])
|
||||||
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value))
|
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value))
|
||||||
const topCrawlerUas = computed(() => countCrawlerUas(crawlers.value).slice(0, 10))
|
const abuseRows = computed(() => formatAbuseRows(data.value?.abuse || [], clients.value, now.value))
|
||||||
|
|
||||||
function flagSvg(code) {
|
|
||||||
return flagSvgs[code?.toUpperCase()] || ''
|
|
||||||
}
|
|
||||||
|
|
||||||
function countryName(code) {
|
|
||||||
if (!code) return ''
|
|
||||||
try {
|
|
||||||
return new Intl.DisplayNames(['en'], { type: 'region' }).of(code.toUpperCase())
|
|
||||||
} catch {
|
|
||||||
return ''
|
|
||||||
}
|
|
||||||
}
|
|
||||||
</script>
|
</script>
|
||||||
|
|
||||||
<template>
|
<template>
|
||||||
@@ -90,8 +126,10 @@ function countryName(code) {
|
|||||||
<p v-else-if="!data" class="loading">loading…</p>
|
<p v-else-if="!data" class="loading">loading…</p>
|
||||||
<template v-else>
|
<template v-else>
|
||||||
<section class="totals">
|
<section class="totals">
|
||||||
<div><strong>{{ visits.length }}</strong> visits</div>
|
<div><strong :title="String(visits.length)">{{ formatCount(visits.length) }}</strong> visits</div>
|
||||||
<div><strong>{{ totalViews }}</strong> page views</div>
|
<div><strong :title="String(totalViews)">{{ formatCount(totalViews) }}</strong> page views</div>
|
||||||
|
<div><strong>{{ readStats.avgMinPerVisit }}</strong> min/visit</div>
|
||||||
|
<div><strong>{{ readStats.avgArticleMedianMin }}</strong> min article read</div>
|
||||||
</section>
|
</section>
|
||||||
|
|
||||||
<VisitorCharts :data="data" :range="range" />
|
<VisitorCharts :data="data" :range="range" />
|
||||||
@@ -103,79 +141,112 @@ function countryName(code) {
|
|||||||
<table class="visit-table">
|
<table class="visit-table">
|
||||||
<thead>
|
<thead>
|
||||||
<tr>
|
<tr>
|
||||||
<th>when</th>
|
|
||||||
<th>trail</th>
|
<th>trail</th>
|
||||||
<th>referer</th>
|
<th>visitor</th>
|
||||||
<th>ip</th>
|
<th class="last-seen">last seen</th>
|
||||||
<th>lang</th>
|
|
||||||
<th>country</th>
|
|
||||||
<th>ua</th>
|
|
||||||
<th>utm</th>
|
|
||||||
</tr>
|
</tr>
|
||||||
</thead>
|
</thead>
|
||||||
<tbody>
|
<tbody>
|
||||||
<tr v-for="(v, i) in visitRows" :key="i">
|
<tr v-for="(v, i) in visitRows" :key="i">
|
||||||
<td class="when">{{ v.when }}</td>
|
|
||||||
<td class="trail">
|
<td class="trail">
|
||||||
<a v-for="(s, si) in v.trail" :key="si"
|
<TrailLink v-if="v.refererStep" :step="v.refererStep" @close="$emit('close')" />
|
||||||
:href="s.path" :title="s.title" @click="emit('close')">
|
<span v-if="v.utm && v.utm !== '—'" class="utm-tag small muted" :title="v.utmTitle">{{ v.utm }}</span>
|
||||||
{{ s.slug }}
|
<TrailLink v-for="(s, si) in v.trail" :key="si" :step="s" @close="$emit('close')" />
|
||||||
</a>
|
|
||||||
</td>
|
</td>
|
||||||
<td>{{ v.referer }}</td>
|
<VisitorCell
|
||||||
<td>
|
:ip="v.ip"
|
||||||
<span class="clickable-ip"
|
:ip-display="v.ipDisplay"
|
||||||
:title="`Click to copy full IP: ${v.ip}`"
|
:ua="v.ua"
|
||||||
@click="copyIp(v.ip)">{{ v.ipDisplay }}</span>
|
:ua-raw="v.uaRaw"
|
||||||
</td>
|
:country="v.country"
|
||||||
<td>{{ v.lang }}</td>
|
:city="v.city"
|
||||||
<td class="country">
|
:lang="v.lang"
|
||||||
<span v-if="flagSvg(v.country)" class="flag" v-html="flagSvg(v.country)" :title="countryName(v.country) || v.country"></span>
|
:lang-display="v.langDisplay"
|
||||||
<template v-else>—</template>
|
:is-host="v.isHost"
|
||||||
</td>
|
/>
|
||||||
<td class="ua" :title="v.uaRaw">{{ v.ua }}</td>
|
<td class="last-seen muted"
|
||||||
<td>{{ v.utm }}</td>
|
:title="v.lastSeenLocal"
|
||||||
|
@click="copyList(v.lastSeenIso, $event)">{{ v.lastSeen }}</td>
|
||||||
</tr>
|
</tr>
|
||||||
</tbody>
|
</tbody>
|
||||||
</table>
|
</table>
|
||||||
</div>
|
</div>
|
||||||
<p v-else class="empty">no visits recorded yet</p>
|
<p v-else class="empty">no visits recorded yet</p>
|
||||||
</section>
|
|
||||||
|
|
||||||
<section>
|
|
||||||
<h2>Crawlers</h2>
|
|
||||||
<div v-if="topCrawlerUas.length" class="crawler-top-uas">
|
|
||||||
<p><strong>top UAs:</strong> {{ formatCounts(topCrawlerUas) }}</p>
|
|
||||||
</div>
|
|
||||||
<div v-if="crawlerRows.length" class="visit-table-wrap">
|
<div v-if="crawlerRows.length" class="visit-table-wrap">
|
||||||
<table class="visit-table">
|
<table class="visit-table">
|
||||||
<thead>
|
<thead>
|
||||||
<tr>
|
<tr>
|
||||||
<th>when</th>
|
<th>pages crawled</th>
|
||||||
<th>entry</th>
|
<th>visitor</th>
|
||||||
<th>ip</th>
|
<th class="last-seen">last seen</th>
|
||||||
<th>ua</th>
|
|
||||||
<th>referer</th>
|
|
||||||
<th>query</th>
|
|
||||||
</tr>
|
</tr>
|
||||||
</thead>
|
</thead>
|
||||||
<tbody>
|
<tbody>
|
||||||
<tr v-for="(c, i) in crawlerRows" :key="i">
|
<tr v-for="(c, i) in crawlerRows" :key="i">
|
||||||
<td class="when">{{ c.when }}</td>
|
<td class="trail">
|
||||||
<td>{{ c.entry }}</td>
|
<TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
|
||||||
<td>
|
|
||||||
<span class="clickable-ip"
|
|
||||||
:title="`Click to copy full IP: ${c.ip}`"
|
|
||||||
@click="copyIp(c.ip)">{{ c.ipDisplay }}</span>
|
|
||||||
</td>
|
</td>
|
||||||
<td class="ua" :title="c.uaRaw">{{ c.ua }}</td>
|
<VisitorCell
|
||||||
<td>{{ c.referer }}</td>
|
:ip="c.ip"
|
||||||
<td>{{ c.query }}</td>
|
:ip-display="c.ipDisplay"
|
||||||
|
:ua="c.ua"
|
||||||
|
:ua-raw="c.uaRaw"
|
||||||
|
:country="c.country"
|
||||||
|
:city="c.city"
|
||||||
|
:lang="c.lang"
|
||||||
|
:lang-display="c.langDisplay"
|
||||||
|
:is-host="c.isHost"
|
||||||
|
/>
|
||||||
|
<td class="last-seen muted"
|
||||||
|
:title="c.lastSeenLocal"
|
||||||
|
@click="copyList(c.lastSeenIso, $event)">{{ c.lastSeen }}</td>
|
||||||
</tr>
|
</tr>
|
||||||
</tbody>
|
</tbody>
|
||||||
</table>
|
</table>
|
||||||
</div>
|
</div>
|
||||||
<p v-else class="empty">no crawler hits recorded yet</p>
|
<p v-else class="empty">no crawler hits recorded yet</p>
|
||||||
|
|
||||||
|
<div v-if="abuseRows.length" class="visit-table-wrap">
|
||||||
|
<table class="visit-table">
|
||||||
|
<thead>
|
||||||
|
<tr>
|
||||||
|
<th>paths abused</th>
|
||||||
|
<th>visitor</th>
|
||||||
|
<th class="last-seen">last seen</th>
|
||||||
|
</tr>
|
||||||
|
</thead>
|
||||||
|
<tbody>
|
||||||
|
<tr v-for="(a, i) in abuseRows" :key="i">
|
||||||
|
<td class="trail abuse-list clickable-list"
|
||||||
|
@click="copyList(a.allPaths, $event)">
|
||||||
|
<div class="abuse-items">
|
||||||
|
<span v-for="(p, pi) in a.paths.slice(0, ABUSE_MAX_LINES)" :key="pi"
|
||||||
|
class="inline-item">
|
||||||
|
<small v-if="p.count > 1" class="muted">{{ formatCount(p.count) }}×</small>{{ p.path }}
|
||||||
|
</span>
|
||||||
|
<small v-if="a.paths.length > ABUSE_MAX_LINES" class="muted">+{{ a.paths.length - ABUSE_MAX_LINES }} more</small>
|
||||||
|
</div>
|
||||||
|
</td>
|
||||||
|
<VisitorCell
|
||||||
|
:ip="a.ip"
|
||||||
|
:ip-display="a.ipDisplay"
|
||||||
|
:ua="a.ua"
|
||||||
|
:ua-raw="a.uaRaw"
|
||||||
|
:country="a.country"
|
||||||
|
:city="a.city"
|
||||||
|
:lang="a.lang"
|
||||||
|
:lang-display="a.langDisplay"
|
||||||
|
:is-host="a.isHost"
|
||||||
|
:variant-count="a.clientCount"
|
||||||
|
/>
|
||||||
|
<td class="last-seen muted"
|
||||||
|
:title="a.lastSeenLocal"
|
||||||
|
@click="copyList(a.lastSeenIso, $event)">{{ a.lastSeen }}</td>
|
||||||
|
</tr>
|
||||||
|
</tbody>
|
||||||
|
</table>
|
||||||
|
</div>
|
||||||
</section>
|
</section>
|
||||||
</template>
|
</template>
|
||||||
</div>
|
</div>
|
||||||
@@ -190,9 +261,10 @@ function countryName(code) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
.analytics-panel {
|
.analytics-panel {
|
||||||
margin: 0 auto;
|
margin: 0;
|
||||||
width: min(60rem, 96vw);
|
width: 100%;
|
||||||
padding: 1.5rem 2rem 4rem;
|
/* Same 1.25rem side spacing as main's article padding. */
|
||||||
|
padding: 1.5rem 1.25rem 4rem;
|
||||||
}
|
}
|
||||||
|
|
||||||
.analytics-panel header {
|
.analytics-panel header {
|
||||||
@@ -215,7 +287,7 @@ function countryName(code) {
|
|||||||
.ranges button {
|
.ranges button {
|
||||||
padding: 0.2rem 0.7rem;
|
padding: 0.2rem 0.7rem;
|
||||||
font: inherit;
|
font: inherit;
|
||||||
font-size: 0.85rem;
|
font-size: 0.9rem;
|
||||||
color: var(--muted);
|
color: var(--muted);
|
||||||
background: none;
|
background: none;
|
||||||
border: 1px solid var(--line);
|
border: 1px solid var(--line);
|
||||||
@@ -250,6 +322,15 @@ function countryName(code) {
|
|||||||
margin-top: 1.8rem;
|
margin-top: 1.8rem;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
.analytics-view a {
|
||||||
|
color: var(--text);
|
||||||
|
text-decoration: none;
|
||||||
|
}
|
||||||
|
.analytics-view a:hover { color: var(--accent); }
|
||||||
|
|
||||||
|
.analytics-view :deep(.muted) { color: var(--muted); }
|
||||||
|
.analytics-view :deep(.small) { font-size: 0.75em; }
|
||||||
|
|
||||||
.totals {
|
.totals {
|
||||||
display: flex;
|
display: flex;
|
||||||
gap: 2rem;
|
gap: 2rem;
|
||||||
@@ -264,8 +345,7 @@ function countryName(code) {
|
|||||||
.visit-table {
|
.visit-table {
|
||||||
width: 100%;
|
width: 100%;
|
||||||
border-collapse: collapse;
|
border-collapse: collapse;
|
||||||
font-family: monospace;
|
font-size: 0.9rem;
|
||||||
font-size: 0.82rem;
|
|
||||||
line-height: 1.3;
|
line-height: 1.3;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -286,9 +366,11 @@ function countryName(code) {
|
|||||||
background: var(--bg, Canvas);
|
background: var(--bg, Canvas);
|
||||||
}
|
}
|
||||||
|
|
||||||
.visit-table .when {
|
.visit-table .last-seen {
|
||||||
|
width: 6rem;
|
||||||
|
text-align: right;
|
||||||
white-space: nowrap;
|
white-space: nowrap;
|
||||||
color: var(--muted);
|
cursor: pointer;
|
||||||
}
|
}
|
||||||
|
|
||||||
.visit-table .trail {
|
.visit-table .trail {
|
||||||
@@ -296,48 +378,74 @@ function countryName(code) {
|
|||||||
overflow-wrap: break-word;
|
overflow-wrap: break-word;
|
||||||
}
|
}
|
||||||
|
|
||||||
.visit-table .trail a {
|
.visit-table .trail a,
|
||||||
color: var(--text);
|
.visit-table .trail-link {
|
||||||
text-decoration: none;
|
display: inline-block;
|
||||||
|
max-width: 8rem;
|
||||||
|
white-space: nowrap;
|
||||||
|
overflow: hidden;
|
||||||
|
text-overflow: ellipsis;
|
||||||
|
vertical-align: bottom;
|
||||||
}
|
}
|
||||||
|
|
||||||
.visit-table .trail a:hover { color: var(--accent); }
|
.visit-table .trail > * + * {
|
||||||
|
|
||||||
.visit-table .trail a + a {
|
|
||||||
margin-left: 0.5rem;
|
margin-left: 0.5rem;
|
||||||
}
|
}
|
||||||
|
|
||||||
.visit-table .clickable-ip {
|
.visit-table .utm-tag {
|
||||||
cursor: pointer;
|
display: inline-block;
|
||||||
text-decoration: underline;
|
max-width: 100%;
|
||||||
text-decoration-style: dotted;
|
padding: 0.05rem 0.4rem;
|
||||||
}
|
border: 1px solid var(--line);
|
||||||
|
border-radius: 0.25rem;
|
||||||
.visit-table .clickable-ip:hover {
|
white-space: nowrap;
|
||||||
color: var(--accent);
|
|
||||||
}
|
|
||||||
|
|
||||||
.visit-table .ua {
|
|
||||||
max-width: 18rem;
|
|
||||||
overflow: hidden;
|
overflow: hidden;
|
||||||
text-overflow: ellipsis;
|
text-overflow: ellipsis;
|
||||||
|
vertical-align: bottom;
|
||||||
|
}
|
||||||
|
|
||||||
|
.visit-table .clickable-list {
|
||||||
|
cursor: pointer;
|
||||||
|
max-width: 22rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.visit-table .abuse-items {
|
||||||
|
display: flex;
|
||||||
|
flex-wrap: wrap;
|
||||||
|
gap: 0.15rem 0.5rem;
|
||||||
|
align-items: baseline;
|
||||||
|
}
|
||||||
|
|
||||||
|
.visit-table .inline-item {
|
||||||
|
max-width: 18rem;
|
||||||
|
min-width: 0;
|
||||||
white-space: nowrap;
|
white-space: nowrap;
|
||||||
}
|
|
||||||
|
|
||||||
.visit-table .country .flag {
|
|
||||||
display: inline-flex;
|
|
||||||
width: 18px;
|
|
||||||
height: 12px;
|
|
||||||
border-radius: 2px;
|
|
||||||
overflow: hidden;
|
overflow: hidden;
|
||||||
border: 1px solid var(--line);
|
text-overflow: ellipsis;
|
||||||
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
|
word-break: keep-all;
|
||||||
|
hyphens: none;
|
||||||
}
|
}
|
||||||
|
|
||||||
.visit-table .country .flag :deep(svg) {
|
.visit-table :deep(.clickable-ip),
|
||||||
width: 100%;
|
.visit-table .clickable-list,
|
||||||
height: 100%;
|
.visit-table .last-seen {
|
||||||
display: block;
|
cursor: pointer;
|
||||||
|
position: relative;
|
||||||
|
}
|
||||||
|
|
||||||
|
.visit-table :deep(.copy-popup) {
|
||||||
|
position: absolute;
|
||||||
|
bottom: calc(100% + 0.25rem);
|
||||||
|
left: 50%;
|
||||||
|
transform: translateX(-50%);
|
||||||
|
padding: 0.15rem 0.4rem;
|
||||||
|
background: var(--text, CanvasText);
|
||||||
|
color: var(--bg, Canvas);
|
||||||
|
border-radius: 0.25rem;
|
||||||
|
font-size: 0.75rem;
|
||||||
|
white-space: nowrap;
|
||||||
|
pointer-events: none;
|
||||||
|
z-index: 10;
|
||||||
}
|
}
|
||||||
|
|
||||||
.crawler-top-uas {
|
.crawler-top-uas {
|
||||||
|
|||||||
+29
-16
@@ -242,30 +242,43 @@ async function saveSettings(opts = {}) {
|
|||||||
async function onThemeChange() {
|
async function onThemeChange() {
|
||||||
await saveSettings()
|
await saveSettings()
|
||||||
// Theme CSS is backend-served at /_themes/{theme}/theme.css in both dev
|
// Theme CSS is backend-served at /_themes/{theme}/theme.css in both dev
|
||||||
// and prod: swap the link in place, then re-render (the theme's default
|
// and prod, but rendered differently: a <link> in dev, an inline <style>
|
||||||
// banner design and the page's stylesheet links may change with it).
|
// in prod. Swap it in place, then re-render (the theme's default banner
|
||||||
let link = document.getElementById('pagerite-theme')
|
// design and the page's stylesheets may change with it).
|
||||||
|
let el = document.getElementById('pagerite-theme')
|
||||||
|
const url = `/_themes/${theme.value}/theme.css`
|
||||||
if (theme.value) {
|
if (theme.value) {
|
||||||
const href = `/_themes/${theme.value}/theme.css`
|
if (el?.tagName === 'STYLE') {
|
||||||
if (link) {
|
el.textContent = await (await fetch(url)).text()
|
||||||
link.href = href
|
} else if (el) {
|
||||||
} else {
|
el.href = url
|
||||||
|
} else if (import.meta.env.DEV) {
|
||||||
// Re-create after "none": keep base < theme < design < custom CSS.
|
// Re-create after "none": keep base < theme < design < custom CSS.
|
||||||
// In dev there is no #pagerite-base link (the base is a
|
// In dev there is no #pagerite-base element (the base is a
|
||||||
// Vite-injected <style>), so anchor to the next sheet instead of
|
// Vite-injected <style>), so anchor to the next sheet instead of
|
||||||
// prepending before the base styles.
|
// prepending before the base styles.
|
||||||
link = document.createElement('link')
|
el = document.createElement('link')
|
||||||
link.rel = 'stylesheet'
|
el.rel = 'stylesheet'
|
||||||
link.id = 'pagerite-theme'
|
el.id = 'pagerite-theme'
|
||||||
link.href = href
|
el.href = url
|
||||||
const before = document.getElementById('pagerite-base')?.nextSibling
|
const before = document.getElementById('pagerite-base')?.nextSibling
|
||||||
?? document.getElementById('pagerite-banner')
|
?? document.getElementById('pagerite-banner')
|
||||||
?? document.getElementById('pagerite-user')
|
?? document.getElementById('pagerite-user')
|
||||||
if (before) before.before(link)
|
if (before) before.before(el)
|
||||||
else document.head.append(link)
|
else document.head.append(el)
|
||||||
|
} else {
|
||||||
|
// Prod: inline <style>, fetched from the backend-served URL.
|
||||||
|
el = document.createElement('style')
|
||||||
|
el.id = 'pagerite-theme'
|
||||||
|
el.textContent = await (await fetch(url)).text()
|
||||||
|
const before = document.getElementById('pagerite-base')?.nextSibling
|
||||||
|
?? document.getElementById('pagerite-banner')
|
||||||
|
?? document.getElementById('pagerite-user')
|
||||||
|
if (before) before.before(el)
|
||||||
|
else document.head.append(el)
|
||||||
}
|
}
|
||||||
} else if (link) {
|
} else if (el) {
|
||||||
link.remove()
|
el.remove()
|
||||||
}
|
}
|
||||||
loadPlain(path.value)
|
loadPlain(path.value)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,22 @@
|
|||||||
|
<script setup>
|
||||||
|
import { formatCount } from './analytics/format.js'
|
||||||
|
|
||||||
|
defineProps({
|
||||||
|
step: { type: Object, required: true },
|
||||||
|
count: { type: Number, default: 0 },
|
||||||
|
})
|
||||||
|
|
||||||
|
defineEmits(['close'])
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<template>
|
||||||
|
<a class="trail-link"
|
||||||
|
:href="step.path"
|
||||||
|
:title="count > 1 ? `${step.title} (${count} hits)` : step.title"
|
||||||
|
:target="step.external ? '_blank' : undefined"
|
||||||
|
:rel="step.external ? 'noopener' : undefined"
|
||||||
|
@click="(e) => { if (!step.external) $emit('close') }">
|
||||||
|
<small v-if="count > 1" class="muted">{{ formatCount(count) }}×</small>
|
||||||
|
<span>{{ step.slug }}</span>
|
||||||
|
</a>
|
||||||
|
</template>
|
||||||
@@ -6,15 +6,18 @@
|
|||||||
* count), so the graph sums the buckets falling inside the selected
|
* count), so the graph sums the buckets falling inside the selected
|
||||||
* range, exactly like the charts and per-page views do.
|
* range, exactly like the charts and per-page views do.
|
||||||
*/
|
*/
|
||||||
import { computed, onBeforeUnmount, shallowRef, watch } from 'vue'
|
import { computed, onBeforeUnmount, onMounted, ref, shallowRef, watch } from 'vue'
|
||||||
import { rangeWindow } from './analytics/time.js'
|
import { rangeWindow, WEEK } from './analytics/time.js'
|
||||||
|
import { formatCount } from './analytics/format.js'
|
||||||
import {
|
import {
|
||||||
TNODE_R,
|
TNODE_W,
|
||||||
|
TNODE_H,
|
||||||
BEAD_R,
|
BEAD_R,
|
||||||
BEAD_SPEED,
|
BEAD_SPEED,
|
||||||
buildTransitionGraph,
|
buildTransitionGraph,
|
||||||
filterTransitionsByRange,
|
filterTransitionsByRange,
|
||||||
filterViewsByRange,
|
filterViewsByRange,
|
||||||
|
filterVisitsByRange,
|
||||||
} from './analytics/transitions.js'
|
} from './analytics/transitions.js'
|
||||||
|
|
||||||
const props = defineProps({
|
const props = defineProps({
|
||||||
@@ -25,6 +28,19 @@ const props = defineProps({
|
|||||||
|
|
||||||
const window = computed(() => rangeWindow(props.range))
|
const window = computed(() => rangeWindow(props.range))
|
||||||
|
|
||||||
|
const visualScale = computed(() => {
|
||||||
|
const { t0, t1 } = window.value
|
||||||
|
if (t0 != null && t1 != null) return WEEK / (t1 - t0)
|
||||||
|
// 'all': scale by the actual data span.
|
||||||
|
const times = new Set()
|
||||||
|
for (const buckets of Object.values(props.data?.views || {})) {
|
||||||
|
for (const k of Object.keys(buckets)) times.add(Date.parse(k))
|
||||||
|
}
|
||||||
|
const arr = [...times]
|
||||||
|
if (arr.length < 2) return 1
|
||||||
|
return WEEK / (Math.max(...arr) - Math.min(...arr))
|
||||||
|
})
|
||||||
|
|
||||||
const filteredData = computed(() => {
|
const filteredData = computed(() => {
|
||||||
if (!props.data) return null
|
if (!props.data) return null
|
||||||
const { t0, t1 } = window.value
|
const { t0, t1 } = window.value
|
||||||
@@ -34,9 +50,13 @@ const filteredData = computed(() => {
|
|||||||
}
|
}
|
||||||
})
|
})
|
||||||
|
|
||||||
|
const filteredVisits = computed(() =>
|
||||||
|
filterVisitsByRange(props.data?.visits, window.value.t0, window.value.t1),
|
||||||
|
)
|
||||||
|
|
||||||
const graph = computed(() =>
|
const graph = computed(() =>
|
||||||
filteredData.value
|
filteredData.value
|
||||||
? buildTransitionGraph(filteredData.value, props.pageTree)
|
? buildTransitionGraph(filteredData.value, props.pageTree, filteredVisits.value, visualScale.value)
|
||||||
: null,
|
: null,
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -47,16 +67,24 @@ const graph = computed(() =>
|
|||||||
const beads = shallowRef([])
|
const beads = shallowRef([])
|
||||||
let rafId = 0
|
let rafId = 0
|
||||||
|
|
||||||
|
const MAX_BEAD_RATE = 120 // upper bound on total beads per second
|
||||||
|
|
||||||
const startBeads = (flows) => {
|
const startBeads = (flows) => {
|
||||||
cancelAnimationFrame(rafId)
|
cancelAnimationFrame(rafId)
|
||||||
beads.value = []
|
beads.value = []
|
||||||
if (!flows?.length) return
|
if (!flows?.length) return
|
||||||
if (matchMedia('(prefers-reduced-motion: reduce)').matches) return
|
if (matchMedia('(prefers-reduced-motion: reduce)').matches) return
|
||||||
|
|
||||||
|
// Cap the total bead emission rate so a busy range cannot spawn enough
|
||||||
|
// beads to kill the page. Existing per-range time scaling is preserved;
|
||||||
|
// this is only a proportional emergency throttle when the limit is hit.
|
||||||
|
const totalRate = flows.reduce((s, f) => s + 1 / f.interval, 0)
|
||||||
|
const scale = totalRate > MAX_BEAD_RATE ? MAX_BEAD_RATE / totalRate : 1
|
||||||
|
|
||||||
const live = [] // { flow, t0 } — one entry per bead in flight
|
const live = [] // { flow, t0 } — one entry per bead in flight
|
||||||
const now = performance.now()
|
const now = performance.now()
|
||||||
const emitters = flows.map((flow) => {
|
const emitters = flows.map((flow) => {
|
||||||
const interval = flow.interval * 1000
|
const interval = (flow.interval / scale) * 1000
|
||||||
// Pre-fill the traversal with evenly spaced beads (random phase), so
|
// Pre-fill the traversal with evenly spaced beads (random phase), so
|
||||||
// the flow appears already running instead of starting empty.
|
// the flow appears already running instead of starting empty.
|
||||||
const phase = Math.random() * interval
|
const phase = Math.random() * interval
|
||||||
@@ -94,32 +122,80 @@ const startBeads = (flows) => {
|
|||||||
|
|
||||||
watch(() => graph.value?.flows, startBeads, { immediate: true })
|
watch(() => graph.value?.flows, startBeads, { immediate: true })
|
||||||
onBeforeUnmount(() => cancelAnimationFrame(rafId))
|
onBeforeUnmount(() => cancelAnimationFrame(rafId))
|
||||||
|
|
||||||
|
// Text in the graph must render at a constant screen size regardless of
|
||||||
|
// how far the enlarged graph's viewBox is scaled down to fit the panel:
|
||||||
|
// measure the unit→pixel ratio and expose it as --u on the svg, which the
|
||||||
|
// font-size rules divide by. Falls back to 1 (raw units) until measured.
|
||||||
|
const svgEl = ref(null)
|
||||||
|
const pxPerUnit = ref(1)
|
||||||
|
let resizeObs = null
|
||||||
|
|
||||||
|
function updateScale() {
|
||||||
|
const el = svgEl.value
|
||||||
|
if (el && el.viewBox.baseVal.width) {
|
||||||
|
pxPerUnit.value = el.clientWidth / el.viewBox.baseVal.width
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
onMounted(() => {
|
||||||
|
resizeObs = new ResizeObserver(updateScale)
|
||||||
|
})
|
||||||
|
watch(svgEl, (el) => {
|
||||||
|
resizeObs?.disconnect()
|
||||||
|
if (el) resizeObs?.observe(el)
|
||||||
|
})
|
||||||
|
watch(() => graph.value?.bounds, updateScale)
|
||||||
|
onBeforeUnmount(() => resizeObs?.disconnect())
|
||||||
|
|
||||||
|
// Font size (px) that fits a label inside the pill width at the current
|
||||||
|
// zoom: ~0.52 em average glyph width, 12 px padding per side, capped.
|
||||||
|
const fitPx = (label) =>
|
||||||
|
Math.min(15, (TNODE_W * pxPerUnit.value - 24) / (0.52 * Math.max(label.length, 1)))
|
||||||
</script>
|
</script>
|
||||||
|
|
||||||
<template>
|
<template>
|
||||||
<section v-if="graph">
|
<section v-if="graph">
|
||||||
<svg class="tmap" :viewBox="`${graph.bounds.x0} ${graph.bounds.y0} ${graph.bounds.x1 - graph.bounds.x0} ${graph.bounds.y1 - graph.bounds.y0}`"
|
<svg ref="svgEl" class="tmap" :style="{ '--u': pxPerUnit }" :viewBox="`${graph.bounds.x0} ${graph.bounds.y0} ${graph.bounds.x1 - graph.bounds.x0} ${graph.bounds.y1 - graph.bounds.y0}`"
|
||||||
role="img" aria-label="map of transitions between pages">
|
role="img" aria-label="map of transitions between pages">
|
||||||
<path v-for="(a, i) in graph.arcs" :key="'a' + i"
|
<path v-for="(a, i) in graph.arcs" :key="'a' + i"
|
||||||
:d="a.d" class="tarc" />
|
:d="a.d" class="tarc" />
|
||||||
<path v-for="(e, i) in graph.edges" :key="'e' + i"
|
<path v-for="(e, i) in graph.edges" :key="'e' + i"
|
||||||
:d="e.d" class="tconn">
|
:d="e.d" :class="['tconn', e.external && 'tconn-exit']">
|
||||||
<title>{{ e.title }}</title>
|
<title>{{ e.title }}</title>
|
||||||
</path>
|
</path>
|
||||||
<circle v-for="(b, i) in beads" :key="'b' + i"
|
<circle v-for="(b, i) in beads" :key="'b' + i"
|
||||||
:cx="b.x" :cy="b.y" :r="BEAD_R" class="tbead" />
|
:cx="b.x" :cy="b.y" :r="BEAD_R" class="tbead" />
|
||||||
<g v-for="(x, i) in graph.extNodes" :key="'x' + i">
|
<g v-for="(x, i) in graph.extNodes" :key="'x' + i">
|
||||||
<circle :cx="x.x" :cy="x.y" :r="x.r" class="txnode">
|
<a v-if="x.href" :href="x.href" target="_blank" rel="noopener">
|
||||||
<title>{{ x.path }}</title>
|
<title>{{ x.path }}</title>
|
||||||
</circle>
|
<rect :x="x.x - TNODE_W/2" :y="x.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2"
|
||||||
<text :x="x.x" :y="x.y + x.r + 11" class="txlabel">{{ x.label }}</text>
|
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
|
||||||
|
<text :x="x.x" :y="x.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle"
|
||||||
|
:style="{ '--slug-px': `${fitPx(x.label)}px` }">{{ x.label }}</text>
|
||||||
|
<text :x="x.x" :y="x.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">{{ formatCount(x.count) }}</text>
|
||||||
|
</a>
|
||||||
|
<g v-else>
|
||||||
|
<title>{{ x.path }}</title>
|
||||||
|
<rect :x="x.x - TNODE_W/2" :y="x.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2"
|
||||||
|
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
|
||||||
|
<text :x="x.x" :y="x.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle"
|
||||||
|
:style="{ '--slug-px': `${fitPx(x.label)}px` }">{{ x.label }}</text>
|
||||||
|
<text :x="x.x" :y="x.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">{{ formatCount(x.count) }}</text>
|
||||||
|
</g>
|
||||||
</g>
|
</g>
|
||||||
<g v-for="n in graph.nodes" :key="n.path">
|
<g v-for="n in graph.nodes" :key="n.path">
|
||||||
<a :href="n.path" :title="n.title">
|
<a :href="n.path">
|
||||||
<circle :cx="n.x" :cy="n.y" :r="TNODE_R" class="tnode" />
|
<title>{{ n.title }}</title>
|
||||||
<text :x="n.x" :y="n.y - 2" class="tnodeslug">{{ n.label }}</text>
|
<rect :x="n.x - TNODE_W/2" :y="n.y - TNODE_H/2" :width="TNODE_W" :height="TNODE_H" :rx="TNODE_H/2" class="tnode" />
|
||||||
<text :x="n.x" :y="n.y + 12" class="tnodecount">{{ n.views }}</text>
|
<text :x="n.x" :y="n.y - TNODE_H*0.16" class="tnodeslug" dominant-baseline="middle"
|
||||||
|
:style="{ '--slug-px': `${fitPx(n.label)}px` }">{{ n.label }}</text>
|
||||||
|
<text :x="n.x" :y="n.y + TNODE_H*0.24" class="tnodecount" dominant-baseline="middle">
|
||||||
|
{{ n.readMin ? `${formatCount(n.views)}×${n.readMin}m` : formatCount(n.views) }}
|
||||||
|
</text>
|
||||||
</a>
|
</a>
|
||||||
|
<text v-if="n.crumb" :x="n.x" :y="n.y - TNODE_H/2 - 10"
|
||||||
|
:class="['tnodepath', n.path === '/' && 'tnodepath-home']">{{ n.crumb }}</text>
|
||||||
</g>
|
</g>
|
||||||
</svg>
|
</svg>
|
||||||
</section>
|
</section>
|
||||||
@@ -130,50 +206,56 @@ onBeforeUnmount(() => cancelAnimationFrame(rafId))
|
|||||||
.tmap {
|
.tmap {
|
||||||
display: block;
|
display: block;
|
||||||
width: 100%;
|
width: 100%;
|
||||||
max-width: 36rem;
|
max-width: 100%;
|
||||||
margin: 0 auto;
|
|
||||||
}
|
}
|
||||||
.tmap .tconn {
|
.tmap .tconn {
|
||||||
fill: var(--accent);
|
fill: var(--accent);
|
||||||
opacity: 0.4; /* uniform, not strength-encoded: width carries that */
|
opacity: 0.4; /* uniform, not strength-encoded: width carries that */
|
||||||
}
|
}
|
||||||
|
.tmap .tconn-exit {
|
||||||
|
fill: var(--text);
|
||||||
|
}
|
||||||
.tmap .tbead {
|
.tmap .tbead {
|
||||||
fill: var(--accent);
|
fill: var(--accent);
|
||||||
opacity: 0.85;
|
opacity: 0.85;
|
||||||
filter: drop-shadow(0 0 2.5px var(--accent));
|
filter: drop-shadow(0 0 2.5px var(--accent));
|
||||||
}
|
}
|
||||||
.tmap .txnode {
|
.tmap .txnode {
|
||||||
fill: var(--bg, Canvas);
|
fill: var(--text);
|
||||||
stroke: var(--muted);
|
stroke: none;
|
||||||
stroke-width: 1;
|
|
||||||
}
|
|
||||||
.tmap .txlabel {
|
|
||||||
fill: var(--muted);
|
|
||||||
font-size: 9px;
|
|
||||||
text-anchor: middle;
|
|
||||||
}
|
}
|
||||||
|
.tmap .txnode-source { fill: var(--text); }
|
||||||
|
.tmap .txnode-exit { fill: var(--text); }
|
||||||
.tmap .tarc {
|
.tmap .tarc {
|
||||||
fill: none;
|
fill: none;
|
||||||
stroke: var(--line);
|
stroke: var(--line);
|
||||||
stroke-width: 1;
|
stroke-width: 1;
|
||||||
}
|
}
|
||||||
.tmap .tnode {
|
.tmap .tnode {
|
||||||
fill: var(--bg, Canvas);
|
fill: var(--accent);
|
||||||
stroke: var(--accent);
|
stroke: none;
|
||||||
stroke-width: 1.5;
|
|
||||||
}
|
}
|
||||||
|
/* Text renders at a constant screen size: --u (set from JS) is the
|
||||||
|
viewBox-unit → pixel ratio of the rendered svg, so dividing by it makes
|
||||||
|
the sizes independent of how far the graph is scaled down. */
|
||||||
.tmap .tnodeslug {
|
.tmap .tnodeslug {
|
||||||
fill: var(--text);
|
fill: var(--bg, Canvas);
|
||||||
font-size: 11px;
|
font-size: calc(var(--slug-px, 15px) / var(--u, 1));
|
||||||
text-anchor: middle;
|
text-anchor: middle;
|
||||||
}
|
}
|
||||||
.tmap a { cursor: pointer; }
|
.tmap a { cursor: pointer; }
|
||||||
.tmap a:hover .tnodeslug { fill: var(--accent); }
|
|
||||||
.tmap .tnodecount {
|
.tmap .tnodecount {
|
||||||
fill: var(--muted);
|
fill: var(--bg, Canvas);
|
||||||
font-size: 10px;
|
opacity: 0.75;
|
||||||
|
font-size: calc(13px / var(--u, 1));
|
||||||
text-anchor: middle;
|
text-anchor: middle;
|
||||||
}
|
}
|
||||||
|
.tmap .tnodepath {
|
||||||
|
fill: var(--muted);
|
||||||
|
font-size: calc(11px / var(--u, 1));
|
||||||
|
text-anchor: middle;
|
||||||
|
}
|
||||||
|
.tmap .tnodepath-home { font-size: calc(17px / var(--u, 1)); }
|
||||||
|
|
||||||
section { margin-top: 1.8rem; }
|
section { margin-top: 1.8rem; }
|
||||||
</style>
|
</style>
|
||||||
|
|||||||
@@ -0,0 +1,156 @@
|
|||||||
|
<script setup>
|
||||||
|
// Visitor metadata cell shared by the recent-visits, crawlers, and abuse tables.
|
||||||
|
// Displays IP/network/host, country flag/city, UA, and language when available.
|
||||||
|
// Clicking the IP copies the full address to the clipboard.
|
||||||
|
// ``variantCount`` overrides the UA line to warn when multiple client
|
||||||
|
// fingerprints share the same IP (e.g. a scanner rotating UAs).
|
||||||
|
import { computed } from 'vue'
|
||||||
|
import * as flagSvgs from 'country-flag-icons/string/3x2'
|
||||||
|
import { copyIp, formatLang } from './analytics/format.js'
|
||||||
|
|
||||||
|
const props = defineProps({
|
||||||
|
ip: { type: String, default: '' },
|
||||||
|
ipDisplay: { type: String, default: '—' },
|
||||||
|
ua: { type: String, default: '' },
|
||||||
|
uaRaw: { type: String, default: '' },
|
||||||
|
country: { type: String, default: '' },
|
||||||
|
city: { type: String, default: '' },
|
||||||
|
lang: { type: String, default: '' },
|
||||||
|
langDisplay: { type: String, default: '' },
|
||||||
|
isHost: { type: Boolean, default: false },
|
||||||
|
variantCount: { type: Number, default: 1 },
|
||||||
|
})
|
||||||
|
|
||||||
|
const hasCountry = computed(() => !!(props.country && props.country !== '—'))
|
||||||
|
const hasCity = computed(() => !!(props.city && props.city !== '—'))
|
||||||
|
const hasLocale = computed(() => hasCountry.value || hasCity.value)
|
||||||
|
const langValue = computed(() => props.langDisplay || formatLang(props.lang))
|
||||||
|
const showLang = computed(() => langValue.value && langValue.value !== '—')
|
||||||
|
|
||||||
|
function flagSvg(code) {
|
||||||
|
return flagSvgs[code?.toUpperCase()] || ''
|
||||||
|
}
|
||||||
|
|
||||||
|
function countryName(code) {
|
||||||
|
if (!code) return ''
|
||||||
|
try {
|
||||||
|
return new Intl.DisplayNames(['en'], { type: 'region' }).of(code.toUpperCase())
|
||||||
|
} catch {
|
||||||
|
return ''
|
||||||
|
}
|
||||||
|
}
|
||||||
|
</script>
|
||||||
|
|
||||||
|
<template>
|
||||||
|
<td class="visitor-cell" :class="{ 'host-cell': isHost }">
|
||||||
|
<div class="visitor-rows">
|
||||||
|
<div class="visitor-row">
|
||||||
|
<div class="locale-line">
|
||||||
|
<span v-if="flagSvg(country)" class="flag" v-html="flagSvg(country)" :title="countryName(country) || country"></span>
|
||||||
|
<template v-if="hasCity"><small class="city-name muted">{{ city }}</small></template>
|
||||||
|
<template v-else-if="!hasLocale">—</template>
|
||||||
|
</div>
|
||||||
|
<div class="ip-line">
|
||||||
|
<span class="clickable-ip small muted"
|
||||||
|
:title="ip"
|
||||||
|
@click="copyIp(ip, $event)">{{ ipDisplay }}</span>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
<div class="visitor-row">
|
||||||
|
<div class="ua-line">
|
||||||
|
<small v-if="variantCount > 1" class="muted variant-hint">{{ variantCount }} client variations</small>
|
||||||
|
<small v-else class="muted" :title="uaRaw">{{ ua || '—' }}</small>
|
||||||
|
</div>
|
||||||
|
<div v-if="showLang && variantCount <= 1" class="locale-lang"><small class="muted">{{ langValue }}</small></div>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
</td>
|
||||||
|
</template>
|
||||||
|
|
||||||
|
<style scoped>
|
||||||
|
.visitor-cell {
|
||||||
|
width: 18em;
|
||||||
|
max-width: 18em;
|
||||||
|
overflow: hidden;
|
||||||
|
text-overflow: ellipsis;
|
||||||
|
vertical-align: top;
|
||||||
|
}
|
||||||
|
|
||||||
|
.visitor-cell.host-cell {
|
||||||
|
text-align: right;
|
||||||
|
}
|
||||||
|
|
||||||
|
.visitor-rows {
|
||||||
|
display: flex;
|
||||||
|
flex-direction: column;
|
||||||
|
gap: 0.15rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.visitor-row {
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
justify-content: space-between;
|
||||||
|
gap: 0.5rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.visitor-row > * {
|
||||||
|
min-width: 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
.locale-line,
|
||||||
|
.ip-line,
|
||||||
|
.ua-line {
|
||||||
|
flex: 1 1 auto;
|
||||||
|
overflow: hidden;
|
||||||
|
text-overflow: ellipsis;
|
||||||
|
white-space: nowrap;
|
||||||
|
}
|
||||||
|
|
||||||
|
.locale-line {
|
||||||
|
text-align: left;
|
||||||
|
display: flex;
|
||||||
|
align-items: center;
|
||||||
|
gap: 0.3rem;
|
||||||
|
}
|
||||||
|
|
||||||
|
.ip-line {
|
||||||
|
text-align: right;
|
||||||
|
}
|
||||||
|
|
||||||
|
.ua-line {
|
||||||
|
text-align: left;
|
||||||
|
}
|
||||||
|
|
||||||
|
.locale-lang {
|
||||||
|
flex: 0 0 auto;
|
||||||
|
overflow: hidden;
|
||||||
|
text-overflow: ellipsis;
|
||||||
|
white-space: nowrap;
|
||||||
|
text-align: right;
|
||||||
|
}
|
||||||
|
|
||||||
|
.city-name {
|
||||||
|
display: inline-block;
|
||||||
|
overflow: hidden;
|
||||||
|
text-overflow: ellipsis;
|
||||||
|
white-space: nowrap;
|
||||||
|
vertical-align: middle;
|
||||||
|
}
|
||||||
|
|
||||||
|
.flag {
|
||||||
|
display: inline-flex;
|
||||||
|
width: 18px;
|
||||||
|
height: 12px;
|
||||||
|
border-radius: 2px;
|
||||||
|
overflow: hidden;
|
||||||
|
border: 1px solid var(--line);
|
||||||
|
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
|
||||||
|
vertical-align: middle;
|
||||||
|
}
|
||||||
|
|
||||||
|
.flag :deep(svg) {
|
||||||
|
width: 100%;
|
||||||
|
height: 100%;
|
||||||
|
display: block;
|
||||||
|
}
|
||||||
|
</style>
|
||||||
@@ -2,10 +2,12 @@
|
|||||||
/**
|
/**
|
||||||
* Visitor and page-view smoothed curves for a single shared time range.
|
* Visitor and page-view smoothed curves for a single shared time range.
|
||||||
*/
|
*/
|
||||||
import { computed } from 'vue'
|
import { computed, onMounted, onUnmounted, ref } from 'vue'
|
||||||
import { makeSeries } from './analytics/time.js'
|
import { makeSeries } from './analytics/time.js'
|
||||||
import { CHART_H, CHART_W, buildChart, fmtY } from './analytics/chart.js'
|
import { CHART_H, CHART_W, buildChart, fmtY } from './analytics/chart.js'
|
||||||
|
|
||||||
|
const DAY_REFRESH_MS = 15000
|
||||||
|
|
||||||
const props = defineProps({
|
const props = defineProps({
|
||||||
data: { type: Object, default: null },
|
data: { type: Object, default: null },
|
||||||
range: { type: String, required: true },
|
range: { type: String, required: true },
|
||||||
@@ -22,33 +24,52 @@ const allViews = computed(() => {
|
|||||||
|
|
||||||
const visitSeries = computed(() => makeSeries(props.data?.site_visits, props.range))
|
const visitSeries = computed(() => makeSeries(props.data?.site_visits, props.range))
|
||||||
const viewSeries = computed(() => makeSeries(allViews.value, props.range))
|
const viewSeries = computed(() => makeSeries(allViews.value, props.range))
|
||||||
const unit = computed(() => (props.range === 'week' ? 'h' : 'day'))
|
|
||||||
|
|
||||||
const visitChart = computed(() => buildChart(visitSeries.value))
|
function freqLabel(unit) {
|
||||||
const viewChart = computed(() => buildChart(viewSeries.value))
|
return unit === '5min' ? '5 min' : unit === 'hour' ? 'hourly' : 'daily'
|
||||||
|
}
|
||||||
|
|
||||||
|
const now = ref(Date.now())
|
||||||
|
let refreshInterval = null
|
||||||
|
onMounted(() => {
|
||||||
|
refreshInterval = setInterval(() => { now.value = Date.now() }, DAY_REFRESH_MS)
|
||||||
|
})
|
||||||
|
onUnmounted(() => {
|
||||||
|
if (refreshInterval) clearInterval(refreshInterval)
|
||||||
|
})
|
||||||
|
|
||||||
|
const visitChart = computed(() => buildChart(visitSeries.value, now.value))
|
||||||
|
const viewChart = computed(() => buildChart(viewSeries.value, now.value))
|
||||||
</script>
|
</script>
|
||||||
|
|
||||||
<template>
|
<template>
|
||||||
<section v-for="c in [
|
<section v-for="c in [
|
||||||
{ ylabel: 'visitors', chart: visitChart, empty: 'no visits recorded yet' },
|
{ ylabel: 'visits', chart: visitChart, empty: 'no visits recorded yet' },
|
||||||
{ ylabel: 'views', chart: viewChart, empty: 'no views recorded yet' },
|
{ ylabel: 'views', chart: viewChart, empty: 'no views recorded yet' },
|
||||||
]" :key="c.ylabel">
|
]" :key="c.ylabel">
|
||||||
<template v-if="c.chart">
|
<template v-if="c.chart">
|
||||||
<div class="chartwrap">
|
<div class="chartwrap">
|
||||||
<div class="plot">
|
<div class="plot">
|
||||||
<div class="plotarea">
|
<div class="plotarea">
|
||||||
<span class="yaxis-label">{{ c.ylabel }}/{{ unit }}</span>
|
<span class="yaxis-label">{{ freqLabel(c.chart.unit) }} {{ c.ylabel }}</span>
|
||||||
<svg class="chart" :viewBox="`0 0 ${CHART_W} ${CHART_H}`"
|
<svg class="chart" :viewBox="`0 0 ${CHART_W} ${CHART_H}`"
|
||||||
preserveAspectRatio="none" role="img" :aria-label="`${c.ylabel} per ${unit}`">
|
preserveAspectRatio="none" role="img" :aria-label="`${freqLabel(c.chart.unit)} ${c.ylabel}`">
|
||||||
<line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value"
|
<line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value"
|
||||||
:x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" />
|
:x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" />
|
||||||
<template v-for="t in c.chart.xticks" :key="'t' + t.x">
|
<template v-for="t in c.chart.xticks" :key="'t' + t.x">
|
||||||
<line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H"
|
<line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H"
|
||||||
class="minor vertical" />
|
class="minor vertical" />
|
||||||
</template>
|
</template>
|
||||||
<template v-for="(s, i) in c.chart.series" :key="i">
|
<template v-if="c.chart.bars">
|
||||||
<path v-if="s.area" :d="s.area" class="area" />
|
<rect v-for="(b, i) in c.chart.bars" :key="'b' + i"
|
||||||
<path :d="s.line" class="line" :style="{ opacity: s.opacity }" />
|
:x="b.x" :y="b.y" :width="b.width" :height="b.height" class="bar" />
|
||||||
|
<path :d="c.chart.skyline" class="line" />
|
||||||
|
</template>
|
||||||
|
<template v-else>
|
||||||
|
<template v-for="(s, i) in c.chart.series" :key="i">
|
||||||
|
<path v-if="s.area" :d="s.area" class="area" />
|
||||||
|
<path :d="s.line" class="line" :style="{ opacity: s.opacity }" />
|
||||||
|
</template>
|
||||||
</template>
|
</template>
|
||||||
<line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5"
|
<line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5"
|
||||||
class="axis" />
|
class="axis" />
|
||||||
@@ -62,7 +83,7 @@ const viewChart = computed(() => buildChart(viewSeries.value))
|
|||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
<div v-if="c.chart.series.length > 1" class="legend">
|
<div v-if="c.chart.series && c.chart.series.length > 1" class="legend">
|
||||||
<span v-for="(s, i) in c.chart.series" :key="i" :style="{ opacity: s.opacity }">
|
<span v-for="(s, i) in c.chart.series" :key="i" :style="{ opacity: s.opacity }">
|
||||||
● {{ s.label }}
|
● {{ s.label }}
|
||||||
</span>
|
</span>
|
||||||
@@ -76,7 +97,7 @@ const viewChart = computed(() => buildChart(viewSeries.value))
|
|||||||
/* The svg is stretched (preserveAspectRatio none), so all text lives in
|
/* The svg is stretched (preserveAspectRatio none), so all text lives in
|
||||||
HTML overlays positioned by the same fractions the geometry uses. */
|
HTML overlays positioned by the same fractions the geometry uses. */
|
||||||
.chartwrap {
|
.chartwrap {
|
||||||
padding-left: 2.2rem; /* y labels */
|
padding-left: 2.8rem; /* y labels */
|
||||||
}
|
}
|
||||||
|
|
||||||
.plot {
|
.plot {
|
||||||
@@ -103,11 +124,11 @@ const viewChart = computed(() => buildChart(viewSeries.value))
|
|||||||
|
|
||||||
.ylab {
|
.ylab {
|
||||||
position: absolute;
|
position: absolute;
|
||||||
left: -2.2rem;
|
left: -2.8rem;
|
||||||
width: 1.9rem;
|
width: 2.6rem;
|
||||||
text-align: right;
|
text-align: right;
|
||||||
transform: translateY(50%);
|
transform: translateY(50%);
|
||||||
font-size: 0.7rem;
|
font-size: 0.75rem;
|
||||||
color: var(--muted);
|
color: var(--muted);
|
||||||
font-variant-numeric: tabular-nums;
|
font-variant-numeric: tabular-nums;
|
||||||
}
|
}
|
||||||
@@ -116,7 +137,7 @@ const viewChart = computed(() => buildChart(viewSeries.value))
|
|||||||
position: absolute;
|
position: absolute;
|
||||||
top: 0.25rem;
|
top: 0.25rem;
|
||||||
transform: translateX(-50%);
|
transform: translateX(-50%);
|
||||||
font-size: 0.7rem;
|
font-size: 0.75rem;
|
||||||
color: var(--muted);
|
color: var(--muted);
|
||||||
white-space: nowrap;
|
white-space: nowrap;
|
||||||
}
|
}
|
||||||
@@ -154,6 +175,11 @@ const viewChart = computed(() => buildChart(viewSeries.value))
|
|||||||
opacity: 0.15;
|
opacity: 0.15;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
.chart .bar {
|
||||||
|
fill: var(--accent);
|
||||||
|
opacity: 0.15;
|
||||||
|
}
|
||||||
|
|
||||||
.chart .line {
|
.chart .line {
|
||||||
fill: none;
|
fill: none;
|
||||||
stroke: var(--accent);
|
stroke: var(--accent);
|
||||||
@@ -176,10 +202,11 @@ const viewChart = computed(() => buildChart(viewSeries.value))
|
|||||||
.yaxis-label {
|
.yaxis-label {
|
||||||
position: absolute;
|
position: absolute;
|
||||||
top: 50%;
|
top: 50%;
|
||||||
left: -2.2rem;
|
left: -2.8rem;
|
||||||
font-size: 0.7rem;
|
font-size: 0.75rem;
|
||||||
color: var(--muted);
|
color: var(--muted);
|
||||||
writing-mode: vertical-rl;
|
writing-mode: vertical-rl;
|
||||||
|
white-space: nowrap;
|
||||||
transform: translateY(-50%) rotate(180deg);
|
transform: translateY(-50%) rotate(180deg);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -1,6 +1,9 @@
|
|||||||
// Analytics page entry: mounts AnalyticsView inside the normal page layout.
|
// Analytics page entry: mounts AnalyticsView inside the normal page layout.
|
||||||
// The backend renders #analytics-app inside #main and links this module for
|
// In production the backend inlines this module into the /_a page (and
|
||||||
// the initial load; pagerite.js also imports it on fetch-navigation to /_a.
|
// pagerite.js re-creates the script element after fetch-navigations there);
|
||||||
|
// in dev pagerite.js imports it from the Vite dev server on demand. Either
|
||||||
|
// way it auto-mounts on #analytics-app when it evaluates, and unmounts when
|
||||||
|
// pagerite.js announces a swap away from /_a.
|
||||||
import { createApp } from 'vue'
|
import { createApp } from 'vue'
|
||||||
import AnalyticsView from './AnalyticsView.vue'
|
import AnalyticsView from './AnalyticsView.vue'
|
||||||
|
|
||||||
@@ -8,9 +11,7 @@ let app = null
|
|||||||
|
|
||||||
export function mount(container) {
|
export function mount(container) {
|
||||||
if (app) return
|
if (app) return
|
||||||
app = createApp(AnalyticsView, {
|
app = createApp(AnalyticsView)
|
||||||
initialRange: new URLSearchParams(location.search).get('range') || 'week',
|
|
||||||
})
|
|
||||||
app.mount(container)
|
app.mount(container)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -19,6 +20,11 @@ export function unmount() {
|
|||||||
app = null
|
app = null
|
||||||
}
|
}
|
||||||
|
|
||||||
// Auto-mount on a normal (non-fetch) page load.
|
// pagerite.js calls this before swapping away from /_a; each evaluation
|
||||||
|
// (the inlined production module evaluates fresh on every visit) replaces
|
||||||
|
// the handle.
|
||||||
|
window.__pageriteAnalyticsUnmount = unmount
|
||||||
|
|
||||||
|
// Auto-mount when the page holding #analytics-app is present.
|
||||||
const container = document.getElementById('analytics-app')
|
const container = document.getElementById('analytics-app')
|
||||||
if (container) mount(container)
|
if (container) mount(container)
|
||||||
|
|||||||
@@ -5,7 +5,8 @@
|
|||||||
* rates (hour on the week view, day on month+).
|
* rates (hour on the week view, day on month+).
|
||||||
*/
|
*/
|
||||||
|
|
||||||
import { DAY, HOUR, WEEK, mondayUTC } from './time.js'
|
import { DAY, HOUR, MIN5, WEEK, mondayUTC } from './time.js'
|
||||||
|
import { formatCount } from './format.js'
|
||||||
|
|
||||||
export const CHART_W = 720
|
export const CHART_W = 720
|
||||||
export const CHART_H = 180
|
export const CHART_H = 180
|
||||||
@@ -14,9 +15,9 @@ export const PAD_TOP = 14 // room above the highest point
|
|||||||
/**
|
/**
|
||||||
* Y always starts at 0; the max is a multiple of a 1-2-5 major step with at
|
* Y always starts at 0; the max is a multiple of a 1-2-5 major step with at
|
||||||
* most 5 intervals, so labeled ticks are always round and evenly divided.
|
* most 5 intervals, so labeled ticks are always round and evenly divided.
|
||||||
* Values are per-unit rates, so small scales are legitimate (a lone visit
|
* A minimum range of 10 keeps tiny near-zero values (e.g. a single visit)
|
||||||
* smoothes to well under 1/unit) — the floor is 1, not 10. Minor lines
|
* from being enlarged to a fractional scale; minor lines subdivide each
|
||||||
* subdivide each major step in five when that yields integers.
|
* major step in five when that yields integers.
|
||||||
*/
|
*/
|
||||||
export function yScale(maxValue) {
|
export function yScale(maxValue) {
|
||||||
let step = 1
|
let step = 1
|
||||||
@@ -27,9 +28,9 @@ export function yScale(maxValue) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
let max = Math.ceil(maxValue / step) * step
|
let max = Math.ceil(maxValue / step) * step
|
||||||
if (max < 1) {
|
if (max < 10) {
|
||||||
max = 1
|
max = 10
|
||||||
step = 0.5
|
step = 2
|
||||||
}
|
}
|
||||||
const minor = step >= 5 && step % 5 === 0 ? step / 5 : null
|
const minor = step >= 5 && step % 5 === 0 ? step / 5 : null
|
||||||
return { max, step, minor }
|
return { max, step, minor }
|
||||||
@@ -206,9 +207,10 @@ export function spline(pts) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
/** Build a full chart model from a series descriptor produced by time.js. */
|
/** Build a full chart model from a series descriptor produced by time.js. */
|
||||||
export function buildChart(input) {
|
export function buildChart(input, now = Date.now()) {
|
||||||
if (!input || !input.series.length) return null
|
if (!input || !input.series.length) return null
|
||||||
const { series, t0, t1, rate, binMinutes, unitMinutes } = input
|
if (input.unit === '5min') return buildDayChart(input, now)
|
||||||
|
const { series, t0, t1, rate, binMinutes, unitMinutes, unit } = input
|
||||||
// Values are per-unit rates (hour on the week view, day on month+); the
|
// Values are per-unit rates (hour on the week view, day on month+); the
|
||||||
// y max is derived from the *smoothed* curves so random single-bucket
|
// y max is derived from the *smoothed* curves so random single-bucket
|
||||||
// spikes don't blow up the scale. Smoothing works on raw counts (its edge
|
// spikes don't blow up the scale. Smoothing works on raw counts (its edge
|
||||||
@@ -283,7 +285,91 @@ export function buildChart(input) {
|
|||||||
x: x(t), left: ((t - t0) / (t1 - t0)) * 100,
|
x: x(t), left: ((t - t0) / (t1 - t0)) * 100,
|
||||||
label: fmtTick(t, t1 - t0), line: true,
|
label: fmtTick(t, t1 - t0), line: true,
|
||||||
}))
|
}))
|
||||||
return { max, majors, minors, series: drawn, xticks }
|
return { max, majors, minors, series: drawn, xticks, unit }
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Day view: 5-minute bars for the last 24 hours. Bars are drawn at raw
|
||||||
|
* counts; the skyline uses a projected full-bucket value for the still-open
|
||||||
|
* final bucket. The y scale is derived from the projected skyline maximum.
|
||||||
|
*/
|
||||||
|
export function buildDayChart(input, now = Date.now()) {
|
||||||
|
const { series, t0, t1 } = input
|
||||||
|
const points = series[0]?.points || []
|
||||||
|
const n = points.length
|
||||||
|
if (!n) return null
|
||||||
|
const bucketMs = (t1 - t0) / n
|
||||||
|
const bucketWidth = CHART_W / n
|
||||||
|
const gap = 0.2
|
||||||
|
const barWidth = Math.max(0.2, bucketWidth - gap)
|
||||||
|
|
||||||
|
const x = (i) => i * bucketWidth + gap / 2
|
||||||
|
const prevRaw = n > 1 ? points[n - 2].count : 0
|
||||||
|
const projected = points.map((p, i) => {
|
||||||
|
if (i !== n - 1) return p.count
|
||||||
|
const bucketStart = t0 + i * bucketMs
|
||||||
|
const elapsed = Math.max(1, Math.min(bucketMs, now - bucketStart))
|
||||||
|
// Blend the observed partial bucket with the previous full bucket:
|
||||||
|
// the longer the current bucket has run, the less we borrow from it.
|
||||||
|
const share = elapsed / bucketMs
|
||||||
|
return p.count + prevRaw * (1 - share)
|
||||||
|
})
|
||||||
|
const highest = Math.max(0, ...projected)
|
||||||
|
const { max, step, minor } = yScale(highest)
|
||||||
|
const y = (v) => PAD_TOP + (1 - Math.max(0, v) / max) * (CHART_H - PAD_TOP)
|
||||||
|
|
||||||
|
const bars = points.map((p, i) => {
|
||||||
|
const bx = x(i)
|
||||||
|
const by = y(p.count)
|
||||||
|
return {
|
||||||
|
x: bx,
|
||||||
|
y: by,
|
||||||
|
width: barWidth,
|
||||||
|
height: CHART_H - by,
|
||||||
|
raw: p.count,
|
||||||
|
projected: projected[i],
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
let skyline = ''
|
||||||
|
for (let i = 0; i < bars.length; i++) {
|
||||||
|
const b = bars[i]
|
||||||
|
const top = y(b.projected)
|
||||||
|
if (i === 0) {
|
||||||
|
skyline += `M${b.x},${top} H${b.x + b.width}`
|
||||||
|
} else {
|
||||||
|
skyline += ` V${top} H${b.x + b.width}`
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const majors = []
|
||||||
|
const minors = []
|
||||||
|
const nMajor = Math.round(max / step)
|
||||||
|
for (let k = 0; k <= nMajor; k++) {
|
||||||
|
const v = k * step
|
||||||
|
majors.push({ value: v, y: y(v), bottom: (1 - PAD_TOP / CHART_H) * (v / max) * 100 })
|
||||||
|
}
|
||||||
|
if (minor) {
|
||||||
|
for (let v = minor; v < max; v += minor) {
|
||||||
|
if (v % step !== 0) minors.push({ y: y(v) })
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const xticks = []
|
||||||
|
const tickStep = 3 * HOUR
|
||||||
|
const firstTick = Math.ceil(t0 / tickStep) * tickStep
|
||||||
|
for (let t = firstTick; t < t1; t += tickStep) {
|
||||||
|
if (t < t0) continue
|
||||||
|
const d = new Date(t)
|
||||||
|
xticks.push({
|
||||||
|
x: ((t - t0) / (t1 - t0)) * CHART_W,
|
||||||
|
left: ((t - t0) / (t1 - t0)) * 100,
|
||||||
|
label: `${String(d.getUTCHours()).padStart(2, '0')}:00`,
|
||||||
|
line: false,
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
return { bars, skyline: skyline.trim(), max, majors, minors, xticks, unit: '5min', series: [] }
|
||||||
}
|
}
|
||||||
|
|
||||||
/** X ticks for year/all: Monday boundaries up to a quarter, UTC month
|
/** X ticks for year/all: Monday boundaries up to a quarter, UTC month
|
||||||
@@ -327,7 +413,7 @@ export function fmtTick(t, span) {
|
|||||||
return d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' })
|
return d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' })
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Y labels: integers when the step allows, one decimal for fractional steps. */
|
/** Y labels use the same compact formatter as text labels. */
|
||||||
export function fmtY(v) {
|
export function fmtY(v) {
|
||||||
return Number.isInteger(v) ? String(v) : v.toFixed(1)
|
return formatCount(v)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -25,11 +25,44 @@ export const hostIP = (ip) => {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Copy the full IP to the clipboard, ignoring failures. */
|
function showCopiedFeedback(el) {
|
||||||
export async function copyIp(ip) {
|
if (!el || typeof document === 'undefined') return
|
||||||
|
const popup = document.createElement('span')
|
||||||
|
popup.textContent = 'Copied!'
|
||||||
|
popup.className = 'copy-popup'
|
||||||
|
popup.style.cssText =
|
||||||
|
'position:absolute;bottom:calc(100% + 0.25rem);left:50%;' +
|
||||||
|
'transform:translateX(-50%);padding:0.15rem 0.4rem;' +
|
||||||
|
'background:var(--text, CanvasText);color:var(--bg, Canvas);' +
|
||||||
|
'border-radius:0.25rem;font-size:0.75rem;white-space:nowrap;' +
|
||||||
|
'pointer-events:none;z-index:10;'
|
||||||
|
el.classList.add('has-copy-popup')
|
||||||
|
el.appendChild(popup)
|
||||||
|
setTimeout(() => {
|
||||||
|
popup.remove()
|
||||||
|
el.classList.remove('has-copy-popup')
|
||||||
|
}, 1200)
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Copy the full IP to the clipboard and show a brief "Copied!" popup. */
|
||||||
|
export async function copyIp(ip, event) {
|
||||||
if (!ip) return
|
if (!ip) return
|
||||||
|
const el = event?.currentTarget
|
||||||
try {
|
try {
|
||||||
await navigator.clipboard.writeText(ip)
|
await navigator.clipboard.writeText(ip)
|
||||||
|
showCopiedFeedback(el)
|
||||||
|
} catch {
|
||||||
|
/* ignore */
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Copy arbitrary text to the clipboard and show a brief "Copied!" popup. */
|
||||||
|
export async function copyList(text, event) {
|
||||||
|
if (!text) return
|
||||||
|
const el = event?.currentTarget
|
||||||
|
try {
|
||||||
|
await navigator.clipboard.writeText(text)
|
||||||
|
showCopiedFeedback(el)
|
||||||
} catch {
|
} catch {
|
||||||
/* ignore */
|
/* ignore */
|
||||||
}
|
}
|
||||||
@@ -44,6 +77,44 @@ export function calcTotalViews(views) {
|
|||||||
return n
|
return n
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Very short reads are navigation/skims, not real reading time.
|
||||||
|
export const MIN_READ_SECONDS = 10
|
||||||
|
|
||||||
|
/** Average minutes per visit and average of per-article median read minutes. */
|
||||||
|
export function calcReadStats(visits) {
|
||||||
|
const perArticle = {}
|
||||||
|
let totalVisitSeconds = 0
|
||||||
|
let visitCount = 0
|
||||||
|
for (const v of visits || []) {
|
||||||
|
const secs = Object.values(v.read || {}).filter((s) => s >= MIN_READ_SECONDS)
|
||||||
|
if (!secs.length) continue
|
||||||
|
visitCount++
|
||||||
|
totalVisitSeconds += secs.reduce((a, b) => a + b, 0)
|
||||||
|
for (const [path, s] of Object.entries(v.read || {})) {
|
||||||
|
if (s >= MIN_READ_SECONDS) {
|
||||||
|
; (perArticle[path] || (perArticle[path] = [])).push(s)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const avgMinPerVisit = visitCount
|
||||||
|
? Math.max(1, Math.round(totalVisitSeconds / visitCount / 60))
|
||||||
|
: 0
|
||||||
|
|
||||||
|
let articleMedianSum = 0
|
||||||
|
const articleCount = Object.keys(perArticle).length
|
||||||
|
for (const arr of Object.values(perArticle)) {
|
||||||
|
arr.sort((a, b) => a - b)
|
||||||
|
const mid = Math.floor(arr.length / 2)
|
||||||
|
const median = arr.length % 2 ? arr[mid] : (arr[mid - 1] + arr[mid]) / 2
|
||||||
|
articleMedianSum += Math.max(MIN_READ_SECONDS, median)
|
||||||
|
}
|
||||||
|
const avgArticleMedianMin = articleCount
|
||||||
|
? Math.max(1, Math.round(articleMedianSum / articleCount / 60))
|
||||||
|
: 0
|
||||||
|
|
||||||
|
return { avgMinPerVisit, avgArticleMedianMin }
|
||||||
|
}
|
||||||
|
|
||||||
/** Build a path -> page title lookup from the site tree. */
|
/** Build a path -> page title lookup from the site tree. */
|
||||||
function buildTitleMap(pageTree) {
|
function buildTitleMap(pageTree) {
|
||||||
const titles = new Map()
|
const titles = new Map()
|
||||||
@@ -59,13 +130,132 @@ function buildTitleMap(pageTree) {
|
|||||||
|
|
||||||
/** Last path segment for display; front page becomes a house icon. */
|
/** Last path segment for display; front page becomes a house icon. */
|
||||||
function slugOf(path) {
|
function slugOf(path) {
|
||||||
return path === '/' ? '🏠' : path.split('/').pop()
|
return path === '/' ? '🏠︎' : path.split('/').pop()
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Host name of an external https origin, with scheme and www. stripped. */
|
||||||
|
function externalSlug(origin) {
|
||||||
|
try {
|
||||||
|
return new URL(origin).host.replace(/^www\./, '')
|
||||||
|
} catch {
|
||||||
|
return origin.replace(/^https?:\/\//, '').replace(/^www\./, '')
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Format one trail step: an internal page or an external https origin. */
|
||||||
|
function stepOf(path, titles) {
|
||||||
|
if (path?.startsWith('/')) {
|
||||||
|
return { path, slug: slugOf(path), title: titles.get(path) || '', external: false, home: path === '/' }
|
||||||
|
}
|
||||||
|
if (path?.startsWith('https://')) {
|
||||||
|
return {
|
||||||
|
path,
|
||||||
|
slug: externalSlug(path),
|
||||||
|
title: 'External site',
|
||||||
|
external: true,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return null
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Human-readable relative timestamp. Adapted from cista-storage: uses
|
||||||
|
* ``Intl.RelativeTimeFormat`` for short intervals and a compact date for
|
||||||
|
* anything older than a week.
|
||||||
|
*/
|
||||||
|
export function formatWhen(ts, now = Date.now()) {
|
||||||
|
const date = new Date(ts)
|
||||||
|
const diff = date.getTime() - now
|
||||||
|
const adiff = Math.abs(diff)
|
||||||
|
const formatter = new Intl.RelativeTimeFormat('en', { numeric: 'auto' })
|
||||||
|
if (adiff <= 5000) return 'now'
|
||||||
|
if (adiff <= 60000) {
|
||||||
|
return formatter
|
||||||
|
.format(Math.round(diff / 1000), 'second')
|
||||||
|
.replace(' ago', '')
|
||||||
|
.replaceAll(' ', '\u202F')
|
||||||
|
}
|
||||||
|
if (adiff <= 3600000) {
|
||||||
|
return formatter
|
||||||
|
.format(Math.round(diff / 60000), 'minute')
|
||||||
|
.replace('utes', '')
|
||||||
|
.replace('ute', '')
|
||||||
|
.replaceAll(' ', '\u202F')
|
||||||
|
}
|
||||||
|
if (adiff <= 86400000) {
|
||||||
|
return formatter
|
||||||
|
.format(Math.round(diff / 3600000), 'hour')
|
||||||
|
.replace('hours', 'h')
|
||||||
|
.replace('hour', 'h')
|
||||||
|
.replaceAll(' ', '\u202F')
|
||||||
|
}
|
||||||
|
if (adiff <= 604800000) {
|
||||||
|
return formatter
|
||||||
|
.format(Math.round(diff / 86400000), 'day')
|
||||||
|
.replaceAll(' ', '\u202F')
|
||||||
|
}
|
||||||
|
let d = date
|
||||||
|
.toLocaleDateString('en-ie', {
|
||||||
|
weekday: 'short',
|
||||||
|
year: 'numeric',
|
||||||
|
month: 'short',
|
||||||
|
day: 'numeric',
|
||||||
|
})
|
||||||
|
.replace('Sept', 'Sep')
|
||||||
|
if (d.length === 14) d = d.replace(' ', ' \u2007')
|
||||||
|
d = d.replaceAll(' ', '\u202F').replace('\u202F', '\u00A0')
|
||||||
|
d = d.slice(0, -4) + d.slice(-2)
|
||||||
|
return d
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Full UTC timestamp for tooltips, e.g. "2026-08-21 00:20:48 UTC". */
|
||||||
|
export function formatWhenTooltip(ts) {
|
||||||
|
return new Date(ts).toISOString().replace('T', ' ').replace('Z', ' UTC')
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Full local timestamp for tooltips, e.g. "21 Aug 2026, 17:38:48". */
|
||||||
|
export function formatWhenLocal(ts) {
|
||||||
|
return new Date(ts).toLocaleString('en-ie', {
|
||||||
|
year: 'numeric',
|
||||||
|
month: 'short',
|
||||||
|
day: 'numeric',
|
||||||
|
hour: '2-digit',
|
||||||
|
minute: '2-digit',
|
||||||
|
second: '2-digit',
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Preserve locale case with the region/country subtag upper-cased. */
|
||||||
|
export function formatLang(value) {
|
||||||
|
if (!value || value === '—') return value
|
||||||
|
const parts = value.split('-')
|
||||||
|
if (parts.length > 1) {
|
||||||
|
parts[parts.length - 1] = parts[parts.length - 1].toUpperCase()
|
||||||
|
}
|
||||||
|
return parts.join('-')
|
||||||
|
}
|
||||||
|
|
||||||
|
/** ISO 8601 UTC timestamp without subseconds, e.g. "2026-08-21T00:20:48Z". */
|
||||||
|
export function formatWhenIso(ts) {
|
||||||
|
return `${new Date(ts).toISOString().split('.')[0]}Z`
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Compact visitor counts: plain below 1k, then 1.2k / 10k / 1.2M.
|
||||||
|
* Truncated, not rounded.
|
||||||
|
*/
|
||||||
|
export function formatCount(n) {
|
||||||
|
if (n < 1000) return String(n)
|
||||||
|
if (n < 10000) return `${Math.trunc(n / 1000)}.${Math.trunc((n % 1000) / 100)}k`
|
||||||
|
if (n < 1_000_000) return `${Math.trunc(n / 1000)}k`
|
||||||
|
return `${Math.trunc(n / 1_000_000)}.${Math.trunc((n % 1_000_000) / 100_000)}M`
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Format recent visits for display, newest first. Each step is a linked slug
|
* Format recent visits for display, newest first. Each step is a linked slug
|
||||||
* pointing to its article; external referers/origins and direct entries are
|
* pointing to its article; external referers/origins are shown as their
|
||||||
* omitted. The link title shows the article heading when known.
|
* domain name with the full origin as the link href. The link title shows the
|
||||||
|
* article heading when known, or "External site" for origins.
|
||||||
*/
|
*/
|
||||||
export function formatRecentVisits(visits, pageTree, limit = 50) {
|
export function formatRecentVisits(visits, pageTree, limit = 50) {
|
||||||
const titles = buildTitleMap(pageTree)
|
const titles = buildTitleMap(pageTree)
|
||||||
@@ -73,13 +263,9 @@ export function formatRecentVisits(visits, pageTree, limit = 50) {
|
|||||||
.reverse()
|
.reverse()
|
||||||
.map((v) => ({
|
.map((v) => ({
|
||||||
when: new Date(v.start).toLocaleString(),
|
when: new Date(v.start).toLocaleString(),
|
||||||
steps: [v.entry, ...(v.trail || [])]
|
steps: [v.referer, v.entry, ...(v.trail || [])]
|
||||||
.filter((p) => p?.startsWith('/'))
|
.map((p) => stepOf(p, titles))
|
||||||
.map((p) => ({
|
.filter(Boolean),
|
||||||
path: p,
|
|
||||||
slug: slugOf(p),
|
|
||||||
title: titles.get(p) || '',
|
|
||||||
})),
|
|
||||||
}))
|
}))
|
||||||
.filter((v) => v.steps.length)
|
.filter((v) => v.steps.length)
|
||||||
.slice(0, limit)
|
.slice(0, limit)
|
||||||
@@ -121,66 +307,228 @@ export function formatCounts(entries) {
|
|||||||
|
|
||||||
/**
|
/**
|
||||||
* Count distinct User-Agent strings among crawler hits, most common first.
|
* Count distinct User-Agent strings among crawler hits, most common first.
|
||||||
* Returns an array of [ua, count] pairs.
|
* Returns an array of [ua, count] pairs. ``clients`` maps client hashes to
|
||||||
|
* client records.
|
||||||
*/
|
*/
|
||||||
export function countCrawlerUas(crawlers) {
|
export function countCrawlerUas(crawlers, clients) {
|
||||||
const counts = {}
|
const counts = {}
|
||||||
for (const c of crawlers || []) {
|
for (const c of crawlers || []) {
|
||||||
const value = c.ua_pretty || c.ua || '(no UA)'
|
const client = (clients || {})[c.client] || {}
|
||||||
|
const value = client.ua_pretty || client.ua || '(no UA)'
|
||||||
counts[value] = (counts[value] || 0) + 1
|
counts[value] = (counts[value] || 0) + 1
|
||||||
}
|
}
|
||||||
return Object.entries(counts).sort((a, b) => b[1] - a[1])
|
return Object.entries(counts).sort((a, b) => b[1] - a[1])
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Format raw crawler hit records as rows for a technical table. Missing
|
* Reduce a reverse-DNS hostname to its right-most components that fit
|
||||||
* values become "—".
|
* within ``limit`` characters. This keeps the meaningful main domain
|
||||||
|
* while avoiding absurdly long subdomains like ``xxx.yyy.zzz...provider.net``.
|
||||||
*/
|
*/
|
||||||
export function formatCrawlerRows(crawlers) {
|
export function mainDomain(host, limit = 24) {
|
||||||
const dash = (s) => (s || '—')
|
if (!host) return host
|
||||||
return [...(crawlers || [])].reverse().map((c) => ({
|
const labels = host.split('.').filter(Boolean)
|
||||||
when: new Date(c.start).toLocaleString(),
|
if (!labels.length) return host
|
||||||
entry: dash(c.entry),
|
const parts = [labels.pop()]
|
||||||
ip: c.ip || '',
|
while (labels.length) {
|
||||||
ipDisplay: c.host || hostIP(c.ip) || c.ip || '—',
|
const next = labels[labels.length - 1]
|
||||||
ua: c.ua_pretty || c.ua || '—',
|
const candidate = `${next}.${parts.join('.')}`
|
||||||
uaRaw: c.ua || '',
|
if (candidate.length > limit) break
|
||||||
referer: dash(c.referer),
|
parts.unshift(labels.pop())
|
||||||
query: dash(c.query),
|
}
|
||||||
}))
|
return parts.join('.')
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Group raw crawler hits by client hash and format each group as a row showing
|
||||||
|
* every internal page that crawler visited. Rows are sorted by total hits,
|
||||||
|
* most active crawler first, rather than by most recent hit.
|
||||||
|
* ``clients`` maps client hashes to client records.
|
||||||
|
*/
|
||||||
|
export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) {
|
||||||
|
const titles = buildTitleMap(pageTree)
|
||||||
|
const groups = new Map()
|
||||||
|
for (const c of crawlers || []) {
|
||||||
|
const client = (clients || {})[c.client] || {}
|
||||||
|
const g = groups.get(c.client) || {
|
||||||
|
clientHash: c.client,
|
||||||
|
client,
|
||||||
|
lastStart: 0,
|
||||||
|
pages: new Map(),
|
||||||
|
}
|
||||||
|
const start = new Date(c.start).getTime()
|
||||||
|
if (start > g.lastStart) g.lastStart = start
|
||||||
|
if (c.entry?.startsWith('/')) {
|
||||||
|
g.pages.set(c.entry, (g.pages.get(c.entry) || 0) + 1)
|
||||||
|
}
|
||||||
|
groups.set(c.client, g)
|
||||||
|
}
|
||||||
|
const totalHits = (g) => {
|
||||||
|
let n = 0
|
||||||
|
for (const c of g.pages.values()) n += c
|
||||||
|
return n
|
||||||
|
}
|
||||||
|
return [...groups.values()]
|
||||||
|
.sort((a, b) => totalHits(b) - totalHits(a) || b.lastStart - a.lastStart)
|
||||||
|
.slice(0, 10)
|
||||||
|
.map((g) => {
|
||||||
|
const client = g.client || {}
|
||||||
|
const host = client.host || ''
|
||||||
|
const isHost = !!host
|
||||||
|
return {
|
||||||
|
lastSeen: formatWhen(g.lastStart, now),
|
||||||
|
lastSeenIso: formatWhenIso(g.lastStart),
|
||||||
|
lastSeenLocal: formatWhenLocal(g.lastStart),
|
||||||
|
pages: [...g.pages.entries()]
|
||||||
|
.sort((a, b) => b[1] - a[1])
|
||||||
|
.map(([path, count]) => ({ ...stepOf(path, titles), count })),
|
||||||
|
ip: client.ip || '',
|
||||||
|
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
|
||||||
|
isHost,
|
||||||
|
ua: client.ua_pretty || client.ua || '—',
|
||||||
|
uaRaw: client.ua || '',
|
||||||
|
lang: client.lang || '—',
|
||||||
|
langDisplay: formatLang(client.lang),
|
||||||
|
country: client.country || '—',
|
||||||
|
city: client.city || '—',
|
||||||
|
total: totalHits(g),
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Group abuse hits by IP and format each group as a row with the full paths
|
||||||
|
* probed. Identical paths are collapsed into one entry with their hit count.
|
||||||
|
* Flagged paths (the ones that triggered abuse classification) are lifted to
|
||||||
|
* the top, followed by other 404s, then document GETs from the abuser. Within
|
||||||
|
* each category paths are sorted by count descending, then earliest first.
|
||||||
|
* Rows are sorted by most recent hit first. Visitor metadata comes from the
|
||||||
|
* latest client hash seen for the IP; ``clientCount`` tells the visitor cell
|
||||||
|
* how many distinct client variations the IP produced. Paths are shown
|
||||||
|
* verbatim (query string included), not resolved against the page tree.
|
||||||
|
* ``clients`` maps client hashes to client records.
|
||||||
|
*/
|
||||||
|
export function formatAbuseRows(abuse, clients, now = Date.now()) {
|
||||||
|
const groups = new Map()
|
||||||
|
for (const a of abuse || []) {
|
||||||
|
const client = (clients || {})[a.client] || {}
|
||||||
|
const ip = client.ip || ''
|
||||||
|
const g = groups.get(ip) || {
|
||||||
|
ip,
|
||||||
|
pathCounts: new Map(),
|
||||||
|
clientHashes: new Set(),
|
||||||
|
lastStart: 0,
|
||||||
|
lastClient: a.client,
|
||||||
|
}
|
||||||
|
const start = new Date(a.start).getTime()
|
||||||
|
if (start > g.lastStart) {
|
||||||
|
g.lastStart = start
|
||||||
|
g.lastClient = a.client
|
||||||
|
}
|
||||||
|
const path = a.path || ''
|
||||||
|
const existing = g.pathCounts.get(path) || {
|
||||||
|
path,
|
||||||
|
count: 0,
|
||||||
|
firstStart: start,
|
||||||
|
flag: a.flag || false,
|
||||||
|
is_404: a.is_404 || false,
|
||||||
|
}
|
||||||
|
existing.count += 1
|
||||||
|
if (start < existing.firstStart) existing.firstStart = start
|
||||||
|
if (a.flag) existing.flag = true
|
||||||
|
if (!a.is_404) existing.is_404 = false
|
||||||
|
g.pathCounts.set(path, existing)
|
||||||
|
g.clientHashes.add(a.client)
|
||||||
|
groups.set(ip, g)
|
||||||
|
}
|
||||||
|
const totalHits = (g) => {
|
||||||
|
let n = 0
|
||||||
|
for (const p of g.pathCounts.values()) n += p.count
|
||||||
|
return n
|
||||||
|
}
|
||||||
|
return [...groups.values()]
|
||||||
|
.sort((a, b) => b.lastStart - a.lastStart)
|
||||||
|
.slice(0, 10)
|
||||||
|
.map((g) => {
|
||||||
|
const pathCategory = (p) => (p.flag ? 0 : p.is_404 ? 1 : 2)
|
||||||
|
const paths = [...g.pathCounts.values()].sort(
|
||||||
|
(a, b) =>
|
||||||
|
pathCategory(a) - pathCategory(b) ||
|
||||||
|
b.count - a.count ||
|
||||||
|
a.firstStart - b.firstStart,
|
||||||
|
)
|
||||||
|
const client = (clients || {})[g.lastClient] || {}
|
||||||
|
const host = client.host || ''
|
||||||
|
const isHost = !!host
|
||||||
|
return {
|
||||||
|
lastSeen: formatWhen(g.lastStart, now),
|
||||||
|
lastSeenIso: formatWhenIso(g.lastStart),
|
||||||
|
lastSeenLocal: formatWhenLocal(g.lastStart),
|
||||||
|
paths: paths.map((p) => ({
|
||||||
|
path: p.path,
|
||||||
|
count: p.count,
|
||||||
|
flag: p.flag,
|
||||||
|
is_404: p.is_404,
|
||||||
|
})),
|
||||||
|
allPaths: paths
|
||||||
|
.map((p) => (p.count > 1 ? `${p.count}× ${p.path}` : p.path))
|
||||||
|
.join('\n'),
|
||||||
|
clientCount: g.clientHashes.size,
|
||||||
|
ip: client.ip || g.ip,
|
||||||
|
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—',
|
||||||
|
isHost,
|
||||||
|
ua: client.ua_pretty || client.ua || '—',
|
||||||
|
uaRaw: client.ua || '',
|
||||||
|
lang: client.lang || '—',
|
||||||
|
langDisplay: formatLang(client.lang),
|
||||||
|
country: client.country || '—',
|
||||||
|
city: client.city || '—',
|
||||||
|
total: totalHits(g),
|
||||||
|
}
|
||||||
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Format raw visit records as rows for a technical table. Returns objects
|
* Format raw visit records as rows for a technical table. Returns objects
|
||||||
* with display strings; missing values become "—". ``trail`` joins page
|
* with display strings; missing values become "—". ``trail`` starts with the
|
||||||
* titles (when known) with " -> ".
|
* external referer (when present), then the entry page and any further internal
|
||||||
|
* pages or external exit origins. Only the 20 most recent visits are shown.
|
||||||
|
* ``clients`` maps client hashes to client records.
|
||||||
*/
|
*/
|
||||||
export function formatVisitRows(visits, pageTree) {
|
export function formatVisitRows(visits, clients, pageTree, now = Date.now()) {
|
||||||
const titles = buildTitleMap(pageTree)
|
const titles = buildTitleMap(pageTree)
|
||||||
return [...(visits || [])].reverse().map((v) => {
|
return [...(visits || [])].reverse().slice(0, 20).map((v) => {
|
||||||
|
const client = (clients || {})[v.client] || {}
|
||||||
const trail = [v.entry, ...(v.trail || [])]
|
const trail = [v.entry, ...(v.trail || [])]
|
||||||
.filter((p) => p?.startsWith('/'))
|
.map((p) => stepOf(p, titles))
|
||||||
.map((p) => ({
|
.filter(Boolean)
|
||||||
path: p,
|
const utmKeys = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content']
|
||||||
slug: slugOf(p),
|
const utmValues = utmKeys.map((k) => (v.utm || {})[k]).filter(Boolean)
|
||||||
title: titles.get(p) || '',
|
const utm = utmValues.length ? utmValues.join(' · ') : ''
|
||||||
}))
|
const utmTitle = Object.entries(v.utm || {})
|
||||||
const utm = Object.entries(v.utm || {})
|
|
||||||
.map(([k, value]) => `${k}=${value}`)
|
.map(([k, value]) => `${k}=${value}`)
|
||||||
.join(', ')
|
.join(', ')
|
||||||
const dash = (s) => (s || '—')
|
const dash = (s) => (s || '—')
|
||||||
|
const host = client.host || ''
|
||||||
|
const isHost = !!host
|
||||||
return {
|
return {
|
||||||
when: new Date(v.start).toLocaleString(),
|
lastSeen: formatWhen(v.start, now),
|
||||||
|
lastSeenIso: formatWhenIso(v.start),
|
||||||
|
lastSeenLocal: formatWhenLocal(v.start),
|
||||||
|
langDisplay: formatLang(client.lang),
|
||||||
trail,
|
trail,
|
||||||
|
refererStep: stepOf(v.referer, titles),
|
||||||
referer: dash(v.referer),
|
referer: dash(v.referer),
|
||||||
ip: v.ip || '',
|
ip: client.ip || '',
|
||||||
ipDisplay: v.host || hostIP(v.ip) || v.ip || '—',
|
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
|
||||||
host: dash(v.host),
|
isHost,
|
||||||
lang: dash(v.lang),
|
lang: dash(client.lang),
|
||||||
country: dash(v.country),
|
country: dash(client.country),
|
||||||
ua: v.ua_pretty || v.ua || '—',
|
city: dash(client.city),
|
||||||
uaRaw: v.ua || '',
|
ua: client.ua_pretty || client.ua || '—',
|
||||||
|
uaRaw: client.ua || '',
|
||||||
utm: utm || '—',
|
utm: utm || '—',
|
||||||
|
utmTitle,
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -13,10 +13,11 @@ export const DAY = 86400e3
|
|||||||
export const WEEK = 7 * DAY
|
export const WEEK = 7 * DAY
|
||||||
|
|
||||||
export const RANGES = {
|
export const RANGES = {
|
||||||
|
day: { label: 'day', span: DAY, bucket: MIN5 },
|
||||||
week: { label: 'week' },
|
week: { label: 'week' },
|
||||||
month: { label: 'month', span: 30 * DAY, bucket: 6 * HOUR },
|
month: { label: 'month', span: 30 * DAY, bucket: 6 * HOUR },
|
||||||
year: { label: 'year', span: 365 * DAY, bucket: DAY },
|
year: { label: 'year', span: 365 * DAY, bucket: DAY },
|
||||||
all: { label: 'all', span: null, bucket: DAY },
|
all: { label: 'all', span: null, bucket: DAY, minSpan: 30 * DAY },
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Monday 00:00 UTC of the week containing t (epoch day 0 was a Thursday). */
|
/** Monday 00:00 UTC of the week containing t (epoch day 0 was a Thursday). */
|
||||||
@@ -89,16 +90,19 @@ export function weeklySeries(buckets) {
|
|||||||
/**
|
/**
|
||||||
* Rolling window for the non-week ranges (x max = now), counts converted
|
* Rolling window for the non-week ranges (x max = now), counts converted
|
||||||
* to per-day rates (the unit the month+ charts are read in).
|
* to per-day rates (the unit the month+ charts are read in).
|
||||||
|
* Ranges without a fixed span use the full data reach, but never less than
|
||||||
|
* their configured minSpan so the chart keeps a readable minimum x scale.
|
||||||
*/
|
*/
|
||||||
export function rollingSeries(buckets, rangeKey) {
|
export function rollingSeries(buckets, rangeKey) {
|
||||||
const raw = rawTimes(buckets)
|
const raw = rawTimes(buckets)
|
||||||
const times = Object.keys(raw).map(Number)
|
const times = Object.keys(raw).map(Number)
|
||||||
if (!times.length) return null
|
if (!times.length) return null
|
||||||
const { span, bucket } = RANGES[rangeKey]
|
const { span, bucket, minSpan = 0 } = RANGES[rangeKey]
|
||||||
const t1 = Math.floor(Date.now() / bucket) * bucket + bucket
|
const t1 = Math.floor(Date.now() / bucket) * bucket + bucket
|
||||||
|
const earliest = Math.floor(Math.min(...times) / bucket) * bucket
|
||||||
const t0 = span != null
|
const t0 = span != null
|
||||||
? t1 - span
|
? t1 - span
|
||||||
: Math.floor(Math.min(...times) / bucket) * bucket
|
: Math.min(earliest, t1 - minSpan)
|
||||||
const points = []
|
const points = []
|
||||||
for (let t = t0; t < t1; t += bucket) {
|
for (let t = t0; t < t1; t += bucket) {
|
||||||
points.push({ t, count: sumRange(raw, t, t + bucket) })
|
points.push({ t, count: sumRange(raw, t, t + bucket) })
|
||||||
@@ -114,11 +118,36 @@ export function rollingSeries(buckets, rangeKey) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Dispatch to weekly or rolling series based on the selected range. */
|
/**
|
||||||
|
* Day view: raw 5-minute bucket counts for the current 24-hour window.
|
||||||
|
* No smoothing or rate conversion is applied; counts are used as-is.
|
||||||
|
*/
|
||||||
|
export function daySeries(buckets) {
|
||||||
|
const raw = rawTimes(buckets)
|
||||||
|
const now = Date.now()
|
||||||
|
const { span, bucket } = RANGES.day
|
||||||
|
const t1 = Math.floor(now / bucket) * bucket + bucket
|
||||||
|
const t0 = t1 - span
|
||||||
|
const points = []
|
||||||
|
for (let t = t0; t < t1; t += bucket) {
|
||||||
|
points.push({ t, count: raw[t] || 0 })
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
series: [{ points, label: '', opacity: 1, area: false }],
|
||||||
|
t0,
|
||||||
|
t1,
|
||||||
|
rate: 1,
|
||||||
|
binMinutes: bucket / 60e3,
|
||||||
|
unitMinutes: bucket / 60e3,
|
||||||
|
unit: '5min',
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Dispatch to daily, weekly or rolling series based on the selected range. */
|
||||||
export function makeSeries(buckets, rangeKey) {
|
export function makeSeries(buckets, rangeKey) {
|
||||||
return rangeKey === 'week'
|
if (rangeKey === 'day') return daySeries(buckets)
|
||||||
? weeklySeries(buckets)
|
if (rangeKey === 'week') return weeklySeries(buckets)
|
||||||
: rollingSeries(buckets, rangeKey)
|
return rollingSeries(buckets, rangeKey)
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -65,6 +65,12 @@
|
|||||||
box-sizing: border-box;
|
box-sizing: border-box;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/* Links never underline — including SVG link text, which the UA stylesheet
|
||||||
|
underlines by default. */
|
||||||
|
a {
|
||||||
|
text-decoration: none;
|
||||||
|
}
|
||||||
|
|
||||||
html {
|
html {
|
||||||
scroll-behavior: smooth;
|
scroll-behavior: smooth;
|
||||||
/* Native-scrollbar fallback styling (JS off or before pagerite.js runs):
|
/* Native-scrollbar fallback styling (JS off or before pagerite.js runs):
|
||||||
@@ -791,8 +797,11 @@ figure:has(img[width]) {
|
|||||||
|
|
||||||
The rules below re-anchor the bleed for the layouts where the article
|
The rules below re-anchor the bleed for the layouts where the article
|
||||||
is not viewport-centered; each just overrides width/margin-inline, and
|
is not viewport-centered; each just overrides width/margin-inline, and
|
||||||
later rules win at equal specificity. */
|
later rules win at equal specificity. The analytics dashboard uses the
|
||||||
figure:has(.wide) {
|
same breakout directly on its container (div.wide — it is the page's
|
||||||
|
whole content, not a figure). */
|
||||||
|
figure:has(.wide),
|
||||||
|
div.wide {
|
||||||
width: 100vw;
|
width: 100vw;
|
||||||
max-width: none;
|
max-width: none;
|
||||||
margin-inline: calc(50% - 50vw);
|
margin-inline: calc(50% - 50vw);
|
||||||
@@ -912,7 +921,9 @@ article h2 {
|
|||||||
padding: 0.5rem 1rem;
|
padding: 0.5rem 1rem;
|
||||||
}
|
}
|
||||||
|
|
||||||
#sidebar ul {
|
/* Only the main level becomes a horizontal wrapping strip; submenus stay
|
||||||
|
vertical blocks attached under their parent item. */
|
||||||
|
#sidebar > ul {
|
||||||
flex-direction: row;
|
flex-direction: row;
|
||||||
flex-wrap: wrap;
|
flex-wrap: wrap;
|
||||||
gap: 0.5rem 1.2rem;
|
gap: 0.5rem 1.2rem;
|
||||||
|
|||||||
+252
-52
@@ -55,6 +55,19 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
let isAdmin = false;
|
let isAdmin = false;
|
||||||
let editorMeta = null;
|
let editorMeta = null;
|
||||||
|
|
||||||
|
// Asset URLs for the on-demand bundles. Dev renders them as
|
||||||
|
// pagerite:* meta tags (Vite dev-server URLs); production inlines all
|
||||||
|
// page assets and carries the on-demand URLs in a JSON script instead.
|
||||||
|
const assets = (() => {
|
||||||
|
const el = document.getElementById("pagerite-assets");
|
||||||
|
if (el) return JSON.parse(el.textContent);
|
||||||
|
const map = {};
|
||||||
|
for (const m of document.querySelectorAll('meta[name^="pagerite:"]')) {
|
||||||
|
map[m.name] = m.content;
|
||||||
|
}
|
||||||
|
return map;
|
||||||
|
})();
|
||||||
|
|
||||||
function makePen(mode) {
|
function makePen(mode) {
|
||||||
const btn = document.createElement("button");
|
const btn = document.createElement("button");
|
||||||
btn.type = "button";
|
btn.type = "button";
|
||||||
@@ -124,11 +137,11 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
}
|
}
|
||||||
|
|
||||||
async function setupAuth() {
|
async function setupAuth() {
|
||||||
const src = document.querySelector('meta[name="pagerite:editor-src"]')?.content;
|
const src = assets["pagerite:editor-src"];
|
||||||
if (!src) { pingEntryOnce(); return; }
|
if (!src) { pingEntryOnce(); return; }
|
||||||
editorMeta = {
|
editorMeta = {
|
||||||
src,
|
src,
|
||||||
css: document.querySelector('meta[name="pagerite:editor-css"]')?.content,
|
css: assets["pagerite:editor-css"],
|
||||||
};
|
};
|
||||||
|
|
||||||
// Detect whether Paskia SSO is available on this site.
|
// Detect whether Paskia SSO is available on this site.
|
||||||
@@ -147,6 +160,26 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
// No auth proxy / dev.
|
// No auth proxy / dev.
|
||||||
}
|
}
|
||||||
|
|
||||||
|
if (isAdmin) {
|
||||||
|
// Teach the backend the site's public origin (used for absolute
|
||||||
|
// social/canonical URLs): unlike request headers, location.origin
|
||||||
|
// reflects the real scheme and host even behind reverse proxies.
|
||||||
|
fetch("/_api/site-url", {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "content-type": "application/json" },
|
||||||
|
body: JSON.stringify({ url: location.origin }),
|
||||||
|
}).catch(() => {});
|
||||||
|
// Warm the cache with the editor bundle: the hashed asset is
|
||||||
|
// immutable, so preloading costs nothing and the pens then open
|
||||||
|
// instantly. The analytics page has no editor.
|
||||||
|
if (currentPath !== "/_a" && !import.meta.env.DEV) {
|
||||||
|
const preload = document.createElement("link");
|
||||||
|
preload.rel = "modulepreload";
|
||||||
|
preload.href = src;
|
||||||
|
document.head.append(preload);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
renderAuthUi();
|
renderAuthUi();
|
||||||
pingEntryOnce();
|
pingEntryOnce();
|
||||||
}
|
}
|
||||||
@@ -230,6 +263,7 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
// buttons; re-add whichever auth UI is appropriate for this session.
|
// buttons; re-add whichever auth UI is appropriate for this session.
|
||||||
renderAuthUi();
|
renderAuthUi();
|
||||||
placeEditPen();
|
placeEditPen();
|
||||||
|
fitNav();
|
||||||
// Multi-column layout only when there is enough text to justify it.
|
// Multi-column layout only when there is enough text to justify it.
|
||||||
// Split the body into columned segments: h1s, h2s and wide figures are
|
// Split the body into columned segments: h1s, h2s and wide figures are
|
||||||
// full-width separators and never go inside columns.
|
// full-width separators and never go inside columns.
|
||||||
@@ -285,14 +319,17 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
// internal link is fetched exactly once, and navigation is served from
|
// internal link is fetched exactly once, and navigation is served from
|
||||||
// memory with no fetch at all. Editor re-renders (swapdoc.loadPlain)
|
// memory with no fetch at all. Editor re-renders (swapdoc.loadPlain)
|
||||||
// announce their fresh copies via pagerite:page-fetched, keeping the
|
// announce their fresh copies via pagerite:page-fetched, keeping the
|
||||||
// cache in sync after edits.
|
// cache in sync after edits. The current page is NOT preloaded: we just
|
||||||
|
// received it as the document (re-fetching would be redundant, and
|
||||||
|
// browser heuristics may send it without if-none-match, defeating the
|
||||||
|
// conditional request); it enters the cache when navigated to.
|
||||||
const pageCache = new Map(); // pathname -> HTML text
|
const pageCache = new Map(); // pathname -> HTML text
|
||||||
addEventListener("pagerite:page-fetched", (ev) => {
|
addEventListener("pagerite:page-fetched", (ev) => {
|
||||||
pageCache.set(new URL(ev.detail.url, location.href).pathname, ev.detail.html);
|
pageCache.set(new URL(ev.detail.url, location.href).pathname, ev.detail.html);
|
||||||
});
|
});
|
||||||
|
|
||||||
function preload() {
|
function preload() {
|
||||||
const urls = new Set([location.pathname]);
|
const urls = new Set();
|
||||||
for (const a of document.querySelectorAll(
|
for (const a of document.querySelectorAll(
|
||||||
'#nav a[href^="/"], #sidebar a[href^="/"], #main a[href^="/"]',
|
'#nav a[href^="/"], #sidebar a[href^="/"], #main a[href^="/"]',
|
||||||
)) {
|
)) {
|
||||||
@@ -300,7 +337,10 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
}
|
}
|
||||||
for (const url of urls) {
|
for (const url of urls) {
|
||||||
if (pageCache.has(url)) continue;
|
if (pageCache.has(url)) continue;
|
||||||
fetch(url)
|
// x-pagerite-preload: idle cache warm-up, not a page view — the
|
||||||
|
// server excludes these GETs from analytics (the ping sent on actual
|
||||||
|
// navigation does the counting).
|
||||||
|
fetch(url, { headers: { "x-pagerite-preload": "1" } })
|
||||||
.then((r) => (r.ok && (r.headers.get("content-type") || "").includes("text/html")
|
.then((r) => (r.ok && (r.headers.get("content-type") || "").includes("text/html")
|
||||||
? r.text() : ""))
|
? r.text() : ""))
|
||||||
.then((html) => { if (html) pageCache.set(url, html); })
|
.then((html) => { if (html) pageCache.set(url, html); })
|
||||||
@@ -338,29 +378,107 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
}
|
}
|
||||||
|
|
||||||
// --- Analytics pings ---------------------------------------------------
|
// --- Analytics pings ---------------------------------------------------
|
||||||
// Fire-and-forget POST /_a {fr, to}: on the initial page load (starts the
|
// Fire-and-forget POST /_a {fr, to, read}: on the initial page load
|
||||||
// visit — the server counts nothing from the document GET alone), for
|
// (starts the visit — the server counts nothing from the document GET
|
||||||
// internal fetch-navigations and for external https exits. Excluded:
|
// alone), for internal fetch-navigations, for external https exits, and
|
||||||
// back/forward (popstate never pings) and everything while we know the
|
// on window close. ``read`` is the active time (ms) spent on ``fr``.
|
||||||
// user is an admin — but only when SSO is actually in use; with no auth
|
// Reading time pauses after 1 minute of inactivity and resumes on the
|
||||||
// (dev/test) "admin" is everyone's state and nothing would be recorded —
|
// next mouse/touch/scroll/keyboard event.
|
||||||
// or has the editor open (admin noise, not visits). The analytics page
|
// Excluded: back/forward (popstate never pings), everything while the
|
||||||
// itself (/_a) is also excluded even though fetch-navigation treats it like
|
// editor is open (body.editing — admin noise, not visits), and the
|
||||||
// a normal article.
|
// analytics page itself (/_a), even though fetch-navigation treats it
|
||||||
|
// like a normal article.
|
||||||
|
// Admins (when SSO is actually in use — with no auth proxy "admin" is
|
||||||
|
// everyone's state) ping normally but with hide=1: the server then
|
||||||
|
// records nothing and scrubs any session the same browser accumulated
|
||||||
|
// before logging in, so admins never show up as visits or crawlers.
|
||||||
// See docs/analytics.md.
|
// See docs/analytics.md.
|
||||||
function ping(to, fr = currentPath) {
|
function ping(to, fr = currentPath, read = 0) {
|
||||||
if ((ssoAvailable && isAdmin) || document.body.classList.contains("editing")
|
if (document.body.classList.contains("editing")) return;
|
||||||
|| to === "/_a" || fr === "/_a") return;
|
if ((to && to === "/_a") || fr === "/_a") return;
|
||||||
|
const hide = ssoAvailable && isAdmin ? 1 : 0;
|
||||||
|
const body = JSON.stringify({
|
||||||
|
fr, to, hide,
|
||||||
|
read: Math.max(0, Math.round(read / 1000)),
|
||||||
|
});
|
||||||
try {
|
try {
|
||||||
fetch("/_a", {
|
fetch("/_a", {
|
||||||
method: "POST",
|
method: "POST",
|
||||||
keepalive: true,
|
keepalive: true,
|
||||||
headers: { "content-type": "application/json" },
|
headers: { "content-type": "application/json" },
|
||||||
body: JSON.stringify({ fr, to }),
|
body,
|
||||||
});
|
});
|
||||||
} catch { /* analytics must never break navigation */ }
|
} catch { /* analytics must never break navigation */ }
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Active reading time for the current page. The clock stops after 1 minute
|
||||||
|
// without activity and restarts on the next mouse/touch/scroll/keyboard
|
||||||
|
// event.
|
||||||
|
const INACTIVE_MS = 60_000;
|
||||||
|
let readStart = performance.now();
|
||||||
|
let readElapsed = 0;
|
||||||
|
let reading = true;
|
||||||
|
let readInactivityTimer = null;
|
||||||
|
let closePingedFor = null;
|
||||||
|
|
||||||
|
function markReadActivity() {
|
||||||
|
if (!reading) {
|
||||||
|
reading = true;
|
||||||
|
readStart = performance.now();
|
||||||
|
}
|
||||||
|
clearTimeout(readInactivityTimer);
|
||||||
|
readInactivityTimer = setTimeout(() => {
|
||||||
|
if (reading) {
|
||||||
|
readElapsed += performance.now() - readStart;
|
||||||
|
reading = false;
|
||||||
|
}
|
||||||
|
}, INACTIVE_MS);
|
||||||
|
}
|
||||||
|
|
||||||
|
function takeReadTime() {
|
||||||
|
if (reading) {
|
||||||
|
readElapsed += performance.now() - readStart;
|
||||||
|
readStart = performance.now();
|
||||||
|
}
|
||||||
|
const ms = Math.max(0, Math.round(readElapsed));
|
||||||
|
readElapsed = 0;
|
||||||
|
return ms;
|
||||||
|
}
|
||||||
|
|
||||||
|
function resetReadTime() {
|
||||||
|
readElapsed = 0;
|
||||||
|
reading = true;
|
||||||
|
readStart = performance.now();
|
||||||
|
clearTimeout(readInactivityTimer);
|
||||||
|
}
|
||||||
|
|
||||||
|
function sendClosePing() {
|
||||||
|
if (closePingedFor === currentPath) return;
|
||||||
|
const read = Math.max(0, Math.round(takeReadTime() / 1000));
|
||||||
|
if (read <= 0) return;
|
||||||
|
const hide = ssoAvailable && isAdmin ? 1 : 0;
|
||||||
|
const body = JSON.stringify({ fr: currentPath, hide, read });
|
||||||
|
const blob = new Blob([body], { type: "application/json" });
|
||||||
|
try {
|
||||||
|
if (navigator.sendBeacon) {
|
||||||
|
navigator.sendBeacon("/_a", blob);
|
||||||
|
} else {
|
||||||
|
fetch("/_a", {
|
||||||
|
method: "POST",
|
||||||
|
keepalive: true,
|
||||||
|
headers: { "content-type": "application/json" },
|
||||||
|
body,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
} catch { /* analytics must never break navigation */ }
|
||||||
|
closePingedFor = currentPath;
|
||||||
|
}
|
||||||
|
|
||||||
|
for (const ev of ["mousemove", "mousedown", "touchstart", "touchmove", "scroll", "keydown"]) {
|
||||||
|
addEventListener(ev, markReadActivity, { passive: true });
|
||||||
|
}
|
||||||
|
addEventListener("pagehide", sendClosePing);
|
||||||
|
|
||||||
// The initial page load pings too — it is what starts the visit and
|
// The initial page load pings too — it is what starts the visit and
|
||||||
// counts the entry page view (the document GET alone records nothing).
|
// counts the entry page view (the document GET alone records nothing).
|
||||||
// Sent once per load, after the auth probes so the admin gate applies;
|
// Sent once per load, after the auth probes so the admin gate applies;
|
||||||
@@ -377,29 +495,39 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
|
|
||||||
// --- Analytics page mount/unmount --------------------------------------
|
// --- Analytics page mount/unmount --------------------------------------
|
||||||
// The analytics page is a normal page whose body is rendered by the server
|
// The analytics page is a normal page whose body is rendered by the server
|
||||||
// but whose content is a Vue app. We load the entry module on demand so the
|
// but whose content is a Vue app. In dev the entry module is imported from
|
||||||
// analytics bundle is only fetched when visiting /_a, and unmount the app
|
// the Vite dev server on demand; in production it is inlined into the /_a
|
||||||
// before swapping away so Vue teardown runs cleanly.
|
// page as script#pagerite-js-analytics, which a fetch-navigation swap does
|
||||||
let analyticsUnmount = null;
|
// not execute — re-create the element so the fresh module auto-mounts on
|
||||||
|
// #analytics-app (see analytics-main.js). The module exposes its unmount
|
||||||
|
// as window.__pageriteAnalyticsUnmount.
|
||||||
function teardownAnalytics() {
|
function teardownAnalytics() {
|
||||||
analyticsUnmount?.();
|
// Remove even the server-rendered script element so a later return to
|
||||||
analyticsUnmount = null;
|
// /_a re-mounts from a fresh copy (the module has torn itself down).
|
||||||
|
document.getElementById("pagerite-js-analytics")?.remove();
|
||||||
|
window.__pageriteAnalyticsUnmount?.();
|
||||||
|
window.__pageriteAnalyticsUnmount = null;
|
||||||
}
|
}
|
||||||
|
|
||||||
async function mountAnalytics(doc) {
|
async function mountAnalytics(doc) {
|
||||||
const src = doc.querySelector('meta[name="pagerite:analytics-src"]')?.content;
|
if (!doc.getElementById("analytics-app")) return;
|
||||||
if (!src) {
|
// Already mounted: on a full /_a load the inline script has run.
|
||||||
teardownAnalytics();
|
if (document.getElementById("pagerite-js-analytics")) return;
|
||||||
|
const inline = doc.getElementById("pagerite-js-analytics");
|
||||||
|
if (inline) {
|
||||||
|
const s = document.createElement("script");
|
||||||
|
for (const a of inline.attributes) s.setAttribute(a.name, a.value);
|
||||||
|
s.textContent = inline.textContent;
|
||||||
|
document.body.append(s);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
try {
|
try {
|
||||||
const mod = await import(/* @vite-ignore */ src);
|
// Dev: the cached module auto-mounts only on its first evaluation,
|
||||||
|
// so call mount() explicitly for repeat visits (it no-ops when the
|
||||||
|
// app is already up).
|
||||||
|
const mod = await import(/* @vite-ignore */ assets["pagerite:analytics-src"]);
|
||||||
const container = document.getElementById("analytics-app");
|
const container = document.getElementById("analytics-app");
|
||||||
if (container) {
|
if (container) mod.mount(container);
|
||||||
mod.mount(container);
|
|
||||||
analyticsUnmount = mod.unmount;
|
|
||||||
}
|
|
||||||
} catch (e) {
|
} catch (e) {
|
||||||
console.error("analytics mount failed:", e);
|
console.error("analytics mount failed:", e);
|
||||||
}
|
}
|
||||||
@@ -425,7 +553,11 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
if (!res.ok || !type.includes("text/html")) throw new Error("not a page");
|
if (!res.ok || !type.includes("text/html")) throw new Error("not a page");
|
||||||
// Reflect any redirect the server issued.
|
// Reflect any redirect the server issued.
|
||||||
if (res.redirected) finalUrl = res.url;
|
if (res.redirected) finalUrl = res.url;
|
||||||
doc = new DOMParser().parseFromString(await res.text(), "text/html");
|
const html = await res.text();
|
||||||
|
// Populate the cache too, so returning here (back/forward, or a
|
||||||
|
// self-link in the nav) is served from memory.
|
||||||
|
pageCache.set(new URL(finalUrl, location.href).pathname, html);
|
||||||
|
doc = new DOMParser().parseFromString(html, "text/html");
|
||||||
} catch {
|
} catch {
|
||||||
location.href = url; // fall back to a normal navigation
|
location.href = url; // fall back to a normal navigation
|
||||||
return false;
|
return false;
|
||||||
@@ -452,26 +584,48 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
} else if (oldSidebar) {
|
} else if (oldSidebar) {
|
||||||
oldSidebar.remove();
|
oldSidebar.remove();
|
||||||
}
|
}
|
||||||
// Site-wide custom CSS lives in <head id="pagerite-user"> and must be
|
// Stylesheets live in <head> with stable ids — links in dev, inline
|
||||||
// kept in sync across fetch-navigations. It is kept last in <head>:
|
// <style> elements in production — and must follow the swap: the
|
||||||
// in dev Vite injects the base stylesheet after the server-rendered
|
// analytics sheet exists on /_a only, and theme/banner/custom CSS
|
||||||
// tag, and equal-specificity :root rules are decided by order.
|
// may have changed since this page was loaded. Diff by id, keeping
|
||||||
const oldUserStyle = document.getElementById("pagerite-user");
|
// the fresh document's order; unchanged sheets keep their elements
|
||||||
const newUserStyle = doc.getElementById("pagerite-user");
|
// so their @keyframes are never torn down. Editor-injected sheets
|
||||||
if (oldUserStyle && newUserStyle) {
|
// (data-pagerite, no id) and Vite's dev styles (no id) are left
|
||||||
oldUserStyle.textContent = newUserStyle.textContent;
|
// alone. Mirrors the head sync in swapdoc.js.
|
||||||
document.head.appendChild(oldUserStyle);
|
const sel = 'link[rel="stylesheet"][id], style[id]';
|
||||||
} else if (newUserStyle) {
|
const fresh = [...doc.head.querySelectorAll(sel)];
|
||||||
document.head.appendChild(document.importNode(newUserStyle, true));
|
const freshIds = new Set(fresh.map((el) => el.id));
|
||||||
} else if (oldUserStyle) {
|
for (const el of [...document.head.querySelectorAll(sel)]) {
|
||||||
oldUserStyle.remove();
|
if (!freshIds.has(el.id)) el.remove();
|
||||||
}
|
}
|
||||||
|
let anchor = null;
|
||||||
|
for (const el of fresh) {
|
||||||
|
const cur = document.getElementById(el.id);
|
||||||
|
if (cur && cur.outerHTML === el.outerHTML) {
|
||||||
|
anchor = cur;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
const imported = document.importNode(el, true);
|
||||||
|
if (cur) cur.replaceWith(imported);
|
||||||
|
else if (anchor) anchor.after(imported);
|
||||||
|
else {
|
||||||
|
const base = document.getElementById("pagerite-base");
|
||||||
|
if (base) base.after(imported);
|
||||||
|
else document.head.append(imported);
|
||||||
|
}
|
||||||
|
anchor = imported;
|
||||||
|
}
|
||||||
|
// Custom CSS must stay last: equal-specificity :root rules (font
|
||||||
|
// variables) are decided by order, and in dev Vite injects the base
|
||||||
|
// stylesheet after the server-rendered tag.
|
||||||
|
const userStyle = document.getElementById("pagerite-user");
|
||||||
|
if (userStyle) document.head.appendChild(userStyle);
|
||||||
document.title = doc.title;
|
document.title = doc.title;
|
||||||
// Banners may contain scripts (canvas etc.), content pages may too.
|
// Banners may contain scripts (canvas etc.), content pages may too.
|
||||||
runScripts(document.getElementById("page-banner"));
|
runScripts(document.getElementById("page-banner"));
|
||||||
runScripts(document.getElementById("main"));
|
runScripts(document.getElementById("main"));
|
||||||
applyEffects();
|
applyEffects();
|
||||||
mountAnalytics(document);
|
mountAnalytics(doc);
|
||||||
};
|
};
|
||||||
// Rotating cube page transition (see the FRAGILE block in pagerite.css);
|
// Rotating cube page transition (see the FRAGILE block in pagerite.css);
|
||||||
// mirrored when navigating back through history. Navigation within the
|
// mirrored when navigating back through history. Navigation within the
|
||||||
@@ -530,9 +684,12 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
if (!a || a.target || a.hasAttribute("download")) return;
|
if (!a || a.target || a.hasAttribute("download")) return;
|
||||||
const url = new URL(a.href, location.href);
|
const url = new URL(a.href, location.href);
|
||||||
if (url.origin !== location.origin) {
|
if (url.origin !== location.origin) {
|
||||||
// External link: the browser navigates; just record the exit (https
|
// External link: the browser navigates; record the full https URL so
|
||||||
// origins only, stripped to the origin part server-side anyway).
|
// different links to the same domain stay distinct in analytics.
|
||||||
if (url.protocol === "https:") ping(url.origin);
|
if (url.protocol === "https:") {
|
||||||
|
closePingedFor = currentPath;
|
||||||
|
ping(url.href, currentPath, takeReadTime());
|
||||||
|
}
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
// Same-page anchor links (footnotes etc.): let the browser handle them
|
// Same-page anchor links (footnotes etc.): let the browser handle them
|
||||||
@@ -544,7 +701,12 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
ev.preventDefault();
|
ev.preventDefault();
|
||||||
// Capture the source now: load() updates currentPath before pinging.
|
// Capture the source now: load() updates currentPath before pinging.
|
||||||
const from = currentPath;
|
const from = currentPath;
|
||||||
load(url).then((ok) => { if (ok) ping(url.pathname, from); });
|
load(url).then((ok) => {
|
||||||
|
if (!ok) return;
|
||||||
|
closePingedFor = null;
|
||||||
|
ping(url.pathname, from, takeReadTime());
|
||||||
|
resetReadTime();
|
||||||
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
addEventListener("popstate", () => {
|
addEventListener("popstate", () => {
|
||||||
@@ -615,6 +777,44 @@ import "overlayscrollbars/overlayscrollbars.css";
|
|||||||
fit();
|
fit();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- Nav condense-to-fit -------------------------------------------------
|
||||||
|
// The top nav stays on one row even on too-narrow screens: first the link
|
||||||
|
// gaps shrink, then the nav's side padding, and only in extreme cases the
|
||||||
|
// font size. #nav is replaced on fetch-navigation swaps, so this re-runs
|
||||||
|
// from applyEffects (fresh elements each time); CSS keeps flex-wrap: wrap
|
||||||
|
// as the no-JS fallback.
|
||||||
|
function fitNav() {
|
||||||
|
const nav = document.getElementById("nav");
|
||||||
|
const ul = nav?.querySelector("ul");
|
||||||
|
if (!ul) return;
|
||||||
|
// Restore the themed defaults before measuring.
|
||||||
|
nav.style.fontSize = "";
|
||||||
|
nav.style.paddingInline = "";
|
||||||
|
ul.style.columnGap = "";
|
||||||
|
ul.style.flexWrap = "nowrap";
|
||||||
|
const overflow = () => ul.scrollWidth - ul.clientWidth;
|
||||||
|
if (overflow() <= 0) return;
|
||||||
|
// 1) shrink the gaps between items (down to a fifth of the themed gap)
|
||||||
|
const gap = parseFloat(getComputedStyle(ul).columnGap) || 0;
|
||||||
|
const joints = Math.max(ul.children.length - 1, 1);
|
||||||
|
if (gap > 0) {
|
||||||
|
ul.style.columnGap = `${Math.max(0.2 * gap, gap - overflow() / joints)}px`;
|
||||||
|
}
|
||||||
|
// 2) shrink the nav's side padding (down to 0.4x)
|
||||||
|
if (overflow() > 0) {
|
||||||
|
const pad = parseFloat(getComputedStyle(nav).paddingInlineStart) || 0;
|
||||||
|
nav.style.paddingInline = `${Math.max(0.4 * pad, pad - overflow() / 2)}px`;
|
||||||
|
}
|
||||||
|
// 3) shrink the font to fit what remains
|
||||||
|
if (overflow() > 0) {
|
||||||
|
const fs = parseFloat(getComputedStyle(nav).fontSize);
|
||||||
|
nav.style.fontSize = `${fs * ul.clientWidth / ul.scrollWidth}px`;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
addEventListener("resize", fitNav);
|
||||||
|
document.fonts?.ready.then(fitNav);
|
||||||
|
|
||||||
setupAuth();
|
setupAuth();
|
||||||
applyEffects();
|
applyEffects();
|
||||||
mountAnalytics(document);
|
mountAnalytics(document);
|
||||||
|
|||||||
+26
-19
@@ -59,32 +59,39 @@ function swapRegions(doc) {
|
|||||||
curUserStyle.remove()
|
curUserStyle.remove()
|
||||||
}
|
}
|
||||||
// Theme and other public stylesheets live in <head>, rendered with stable
|
// Theme and other public stylesheets live in <head>, rendered with stable
|
||||||
// ids by the backend; sync them positionally so the custom CSS (rendered
|
// ids by the backend (links in dev, inline <style> elements in prod);
|
||||||
// last) always keeps winning by order. Diff-based: unchanged sheets keep
|
// sync them positionally so the custom CSS (rendered last) always keeps
|
||||||
// their elements, so their @keyframes are never torn down (re-creating
|
// winning by order. Diff-based: unchanged sheets keep their elements, so
|
||||||
// keyframes would replay the editor's slide-in animation).
|
// their @keyframes are never torn down (re-creating keyframes would
|
||||||
const freshLinks = [...doc.head.querySelectorAll('link[rel="stylesheet"]')]
|
// replay the editor's slide-in animation).
|
||||||
const freshIds = new Set(freshLinks.map((l) => l.id))
|
const sel = 'link[rel="stylesheet"][id], style[id]'
|
||||||
for (const link of [...document.head.querySelectorAll('link[rel="stylesheet"]')]) {
|
const freshEls = [...doc.head.querySelectorAll(sel)]
|
||||||
if (!link.dataset.pagerite && !freshIds.has(link.id)) link.remove()
|
const freshIds = new Set(freshEls.map((el) => el.id))
|
||||||
|
for (const el of [...document.head.querySelectorAll(sel)]) {
|
||||||
|
if (!freshIds.has(el.id)) el.remove()
|
||||||
}
|
}
|
||||||
// Insert missing sheets in the fresh document's order, each right after
|
// Insert missing sheets in the fresh document's order, each right after
|
||||||
// its predecessor's element. The first sheet rendered is always the base
|
// its predecessor's element. The first sheet rendered is always the base
|
||||||
// CSS, so its link doubles as the fallback anchor when nothing matched yet
|
// CSS, so its element doubles as the fallback anchor when nothing matched
|
||||||
// (e.g. no theme was selected before and the position is otherwise lost).
|
// yet (e.g. no theme was selected before and the position is otherwise
|
||||||
|
// lost).
|
||||||
let anchor = null
|
let anchor = null
|
||||||
for (const link of freshLinks) {
|
for (const el of freshEls) {
|
||||||
const cur = link.id && document.getElementById(link.id)
|
const cur = el.id && document.getElementById(el.id)
|
||||||
if (cur && cur.href === link.href) {
|
if (cur && cur.outerHTML === el.outerHTML) {
|
||||||
anchor = cur
|
anchor = cur
|
||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
const el = document.importNode(link, true)
|
const imported = document.importNode(el, true)
|
||||||
// Same id, new URL (theme switch): replace in place, keeping position.
|
// Same id, new content (theme switch): replace in place, keeping position.
|
||||||
if (cur) cur.replaceWith(el)
|
if (cur) cur.replaceWith(imported)
|
||||||
else if (anchor) anchor.after(el)
|
else if (anchor) anchor.after(imported)
|
||||||
else document.getElementById('pagerite-base')?.after(el) ?? document.head.append(el)
|
else {
|
||||||
anchor = el
|
const base = document.getElementById('pagerite-base')
|
||||||
|
if (base) base.after(imported)
|
||||||
|
else document.head.append(imported)
|
||||||
|
}
|
||||||
|
anchor = imported
|
||||||
}
|
}
|
||||||
// The editor keeps its own title while open; only inherit the server title
|
// The editor keeps its own title while open; only inherit the server title
|
||||||
// when navigating outside the editor (e.g. fetch-navigation swaps).
|
// when navigating outside the editor (e.g. fetch-navigation swaps).
|
||||||
|
|||||||
@@ -11,7 +11,7 @@
|
|||||||
*/
|
*/
|
||||||
|
|
||||||
export default function fastapiVue({ paths = ["/api"] } = {}) {
|
export default function fastapiVue({ paths = ["/api"] } = {}) {
|
||||||
const backendUrl = process.env.PAGERITE_BACKEND_URL || "http://localhost:3200"
|
const backendUrl = process.env.PAGERITE_BACKEND_URL || "http://localhost:8210"
|
||||||
|
|
||||||
// Build proxy configuration for each path
|
// Build proxy configuration for each path
|
||||||
const proxy = {}
|
const proxy = {}
|
||||||
|
|||||||
@@ -7,11 +7,10 @@ import vueDevTools from 'vite-plugin-vue-devtools'
|
|||||||
|
|
||||||
const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:3200'
|
const backendUrl = process.env.PAGERITE_BACKEND_URL || 'http://localhost:3200'
|
||||||
|
|
||||||
// Proxy content pages (/slug, /path/to/slug) to the FastAPI backend in dev.
|
// Proxy everything except Vite's own dev-time paths and the backend machinery
|
||||||
// Excludes Vite internals (/@..., /src, /node_modules, /__...) and the
|
// to the FastAPI backend in dev. /_api, /_f, /_themes and /_a are handled by
|
||||||
// backend's /_ prefix. /_api, /_f, /_themes and the /_a analytics ping are
|
// the fastapi-vue plugin, and /@..., /src, /node_modules, /__... stay with Vite.
|
||||||
// handled by the fastapi-vue plugin.
|
const CONTENT_PROXY = '^(?!/_|/@|/src|/node_modules|/__).*$'
|
||||||
const CONTENT_PROXY = '^\\/(?!_|@|src|node_modules|__)(?:[^./?]+(?:\\/[^./?]+)*)?(?:\\?.*)?$'
|
|
||||||
|
|
||||||
// https://vite.dev/config/
|
// https://vite.dev/config/
|
||||||
export default defineConfig({
|
export default defineConfig({
|
||||||
|
|||||||
+69
-2
@@ -1,14 +1,74 @@
|
|||||||
# auto-upgrade@fastapi-vue-setup - remove this if you modify this file
|
|
||||||
"""Command-line entry point for running the backend server."""
|
"""Command-line entry point for running the backend server."""
|
||||||
|
|
||||||
import argparse
|
import argparse
|
||||||
|
import gzip
|
||||||
import os
|
import os
|
||||||
|
import sys
|
||||||
|
from datetime import date
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import httpx
|
||||||
from fastapi_vue import server
|
from fastapi_vue import server
|
||||||
|
|
||||||
DEFAULT_PORT = 3100
|
DEFAULT_PORT = 8100
|
||||||
DEVMODE = os.getenv("PAGERITE_DEV") == "1"
|
DEVMODE = os.getenv("PAGERITE_DEV") == "1"
|
||||||
|
|
||||||
|
# Repository root (pagerite/__main__.py -> ..), where the MMDB lives.
|
||||||
|
_REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||||
|
|
||||||
|
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
|
||||||
|
|
||||||
|
|
||||||
|
def _download_dbip() -> None:
|
||||||
|
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
|
||||||
|
today = date.today()
|
||||||
|
months = [f"{today:%Y-%m}"]
|
||||||
|
# The current month's file may not be published yet; fall back to last month.
|
||||||
|
prev = (today.replace(day=1) - date.resolution).replace(day=1)
|
||||||
|
months.append(f"{prev:%Y-%m}")
|
||||||
|
|
||||||
|
existing = sorted(
|
||||||
|
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
|
||||||
|
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
|
||||||
|
)
|
||||||
|
if existing and existing[-1] >= months[0]:
|
||||||
|
print(f"pagerite: DB-IP database is current ({existing[-1]}), skipping download")
|
||||||
|
return
|
||||||
|
|
||||||
|
for month in months:
|
||||||
|
url = DBIP_URL.format(month=month)
|
||||||
|
target = _REPO_ROOT / f"dbip-city-lite-{month}.mmdb.gz"
|
||||||
|
tmp = target.with_suffix(".mmdb.gz.tmp")
|
||||||
|
print(f"pagerite: downloading {url}")
|
||||||
|
try:
|
||||||
|
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
|
||||||
|
if r.status_code == 404:
|
||||||
|
continue
|
||||||
|
r.raise_for_status()
|
||||||
|
with open(tmp, "wb") as f:
|
||||||
|
for chunk in r.iter_bytes():
|
||||||
|
f.write(chunk)
|
||||||
|
except httpx.HTTPError as e:
|
||||||
|
print(f"pagerite: DB-IP download failed: {e}", file=sys.stderr)
|
||||||
|
tmp.unlink(missing_ok=True)
|
||||||
|
continue
|
||||||
|
# Verify it is actually gzip data before installing it.
|
||||||
|
try:
|
||||||
|
with gzip.open(tmp, "rb") as f:
|
||||||
|
f.read(1)
|
||||||
|
except OSError:
|
||||||
|
print(f"pagerite: DB-IP download for {month} was not valid gzip", file=sys.stderr)
|
||||||
|
tmp.unlink(missing_ok=True)
|
||||||
|
continue
|
||||||
|
os.replace(tmp, target)
|
||||||
|
# Drop older databases so the app never picks up a stale one.
|
||||||
|
for old in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*"):
|
||||||
|
if old.name != target.name:
|
||||||
|
old.unlink()
|
||||||
|
print(f"pagerite: DB-IP database updated to {target.name}")
|
||||||
|
return
|
||||||
|
print("pagerite: could not download a DB-IP database", file=sys.stderr)
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
def main() -> None:
|
||||||
"""Run the backend server with optional arguments."""
|
"""Run the backend server with optional arguments."""
|
||||||
@@ -19,7 +79,14 @@ def main() -> None:
|
|||||||
action="append",
|
action="append",
|
||||||
help=(f"Endpoint (default: localhost:{DEFAULT_PORT})."),
|
help=(f"Endpoint (default: localhost:{DEFAULT_PORT})."),
|
||||||
)
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--dbip",
|
||||||
|
action="store_true",
|
||||||
|
help="Download/update the DB-IP city lite database before starting.",
|
||||||
|
)
|
||||||
args = parser.parse_args()
|
args = parser.parse_args()
|
||||||
|
if args.dbip:
|
||||||
|
_download_dbip()
|
||||||
dev = {"reload": True, "reload_dirs": ["pagerite"]} if DEVMODE else {}
|
dev = {"reload": True, "reload_dirs": ["pagerite"]} if DEVMODE else {}
|
||||||
server.run(
|
server.run(
|
||||||
"pagerite.app:app",
|
"pagerite.app:app",
|
||||||
|
|||||||
+467
-100
@@ -5,22 +5,37 @@ ping on page load starts a visit, later pings extend it, and pings with no
|
|||||||
known session start a fresh one (missing data, not dropped). The document
|
known session start a fresh one (missing data, not dropped). The document
|
||||||
GET handler stashes the entry referer (external https origin) and any
|
GET handler stashes the entry referer (external https origin) and any
|
||||||
utm_* query parameters in in-memory IP tables, consumed when the ping
|
utm_* query parameters in in-memory IP tables, consumed when the ping
|
||||||
starts the visit; nothing is counted without a ping (bots and admin
|
starts the visit; nothing is counted without a ping (plain bots that only
|
||||||
browsing stay invisible). The session map is in-memory only. The visitor
|
fetch documents end up in the crawler list). JS-running crawlers
|
||||||
IP and, when available, its reverse-DNS host name are stored on the visit
|
(Googlebot, GoogleOther, Applebot, ...) do ping, but their UA gives them
|
||||||
record itself.
|
away (``_is_bot_ua``) and their pings are ignored, so they land in the
|
||||||
|
crawler list too. Idle-time link preloads from pagerite.js carry an
|
||||||
|
``x-pagerite-preload`` header and are not tracked at all — the ping sent
|
||||||
|
when the user actually navigates does the counting.
|
||||||
|
Admin clients ping with ``hide=1``, which records nothing and removes any
|
||||||
|
visit the session accumulated before logging in. Scanner telltale 404s
|
||||||
|
(dotpaths, *.php) classify the source IP as abuse; its hits — including
|
||||||
|
earlier crawler hits — are moved to the abuse list, which the viewer
|
||||||
|
groups by IP with full request paths. Client metadata (IP, UA, language,
|
||||||
|
country/city, host) is stored once per unique client hash and referenced
|
||||||
|
from visits, crawler hits and abuse hits. The session map is in-memory
|
||||||
|
only.
|
||||||
|
|
||||||
Data is a msgspec Struct JSON-dumped to its own file (not the kanta db),
|
Data is a msgspec Struct JSON-dumped to its own file (not the kanta db),
|
||||||
rewritten atomically on every recorded event.
|
rewritten atomically on every recorded event.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
|
import ipaddress
|
||||||
import os
|
import os
|
||||||
import re
|
import re
|
||||||
import tempfile
|
import tempfile
|
||||||
|
from collections.abc import Callable
|
||||||
|
from contextlib import suppress
|
||||||
from datetime import UTC, datetime, timedelta
|
from datetime import UTC, datetime, timedelta
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from urllib.parse import parse_qs, urlparse
|
from urllib.parse import parse_qs, urlparse
|
||||||
|
|
||||||
|
import blake3
|
||||||
import msgspec
|
import msgspec
|
||||||
from ua_parser import parse
|
from ua_parser import parse
|
||||||
|
|
||||||
@@ -39,7 +54,10 @@ def _compact_user_agent(ua: str) -> str:
|
|||||||
dev = r.device.family if r.device else None
|
dev = r.device.family if r.device else None
|
||||||
if browser in (None, "Other") and os_name in (None, "Other"):
|
if browser in (None, "Other") and os_name in (None, "Other"):
|
||||||
return ua
|
return ua
|
||||||
browser = browser if browser and browser != "Other" else ""
|
if browser and browser != "Other":
|
||||||
|
browser = browser.split()[0]
|
||||||
|
else:
|
||||||
|
browser = ""
|
||||||
os_name = os_name if os_name and os_name != "Other" else ""
|
os_name = os_name if os_name and os_name != "Other" else ""
|
||||||
if dev in (None, "Other") or dev == browser:
|
if dev in (None, "Other") or dev == browser:
|
||||||
dev = ""
|
dev = ""
|
||||||
@@ -47,50 +65,90 @@ def _compact_user_agent(ua: str) -> str:
|
|||||||
return " ".join(p for p in parts if p).strip()
|
return " ".join(p for p in parts if p).strip()
|
||||||
|
|
||||||
|
|
||||||
|
class Client(msgspec.Struct, omit_defaults=True):
|
||||||
|
"""Client metadata shared by visits, crawler hits and abuse hits.
|
||||||
|
|
||||||
|
Identified by a 6-byte blake3 hash of the IPv4 address or IPv6 /64
|
||||||
|
network, the full User-Agent string and the extracted language tag.
|
||||||
|
Country/city/host are filled in asynchronously after the first event.
|
||||||
|
"""
|
||||||
|
|
||||||
|
#: Visitor IP address (first X-Forwarded-For hop or direct peer).
|
||||||
|
ip: str = ""
|
||||||
|
#: Reverse-DNS host name for ``ip`` when resolvable, else "".
|
||||||
|
host: str = ""
|
||||||
|
#: First Accept-Language tag, lowercased (e.g. "en-us").
|
||||||
|
lang: str = ""
|
||||||
|
#: Two-letter country code from the DB-IP geoip lookup, or "".
|
||||||
|
country: str = ""
|
||||||
|
#: City name from the DB-IP geoip lookup, or "".
|
||||||
|
city: str = ""
|
||||||
|
#: Raw User-Agent header.
|
||||||
|
ua: str = ""
|
||||||
|
#: Compact display form of ``ua`` (browser/OS/device) when parsable.
|
||||||
|
ua_pretty: str = ""
|
||||||
|
|
||||||
|
|
||||||
class Visit(msgspec.Struct, omit_defaults=True):
|
class Visit(msgspec.Struct, omit_defaults=True):
|
||||||
"""One visit: the initial-load data plus everything seen afterwards.
|
"""One visit: the initial-load data plus everything seen afterwards.
|
||||||
|
|
||||||
``trail`` holds page paths and external exit origins in first-seen
|
``trail`` holds page paths and external exit URLs in first-seen
|
||||||
order; re-visiting an already seen page does not append. The entry
|
order; re-visiting an already seen page does not append. The entry
|
||||||
page itself is in ``entry``, not in the trail.
|
page itself is in ``entry``, not in the trail. Client metadata is
|
||||||
|
held in ``Analytics.clients`` keyed by ``client``.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
start: datetime
|
start: datetime
|
||||||
entry: str
|
entry: str
|
||||||
#: External https origin of the initial load, "" for direct visits.
|
#: External https origin of the initial load, "" for direct visits.
|
||||||
referer: str = ""
|
referer: str = ""
|
||||||
#: Visitor IP address (first X-Forwarded-For hop or direct peer).
|
#: 6-byte blake3 hash referencing ``Analytics.clients``.
|
||||||
ip: str = ""
|
client: bytes = b""
|
||||||
#: Reverse-DNS host name for ``ip`` when resolvable, else "".
|
|
||||||
host: str = ""
|
|
||||||
trail: list[str] = []
|
trail: list[str] = []
|
||||||
#: First Accept-Language tag, lowercased (e.g. "en-us").
|
|
||||||
lang: str = ""
|
|
||||||
#: Two-letter region subtag derived from ``lang`` (e.g. "US"), or "".
|
|
||||||
country: str = ""
|
|
||||||
#: Raw User-Agent header from the initial ping.
|
|
||||||
ua: str = ""
|
|
||||||
#: Compact display form of ``ua`` (browser/OS/device) when parsable.
|
|
||||||
ua_pretty: str = ""
|
|
||||||
#: UTM query parameters from the landing URL, keyed by parameter name.
|
#: UTM query parameters from the landing URL, keyed by parameter name.
|
||||||
utm: dict[str, str] = {}
|
utm: dict[str, str] = {}
|
||||||
|
#: Active reading time per path (seconds), keyed by path.
|
||||||
|
read: dict[str, int] = {}
|
||||||
|
|
||||||
|
|
||||||
class CrawlerHit(msgspec.Struct, omit_defaults=True):
|
class CrawlerHit(msgspec.Struct, omit_defaults=True):
|
||||||
"""A document GET that was never followed by an analytics ping."""
|
"""A document GET that was never followed by an analytics ping.
|
||||||
|
|
||||||
|
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
|
||||||
|
"""
|
||||||
|
|
||||||
start: datetime
|
start: datetime
|
||||||
entry: str
|
entry: str
|
||||||
ip: str = ""
|
#: 6-byte blake3 hash referencing ``Analytics.clients``.
|
||||||
ua: str = ""
|
client: bytes = b""
|
||||||
#: Compact display form of ``ua`` when parsable.
|
|
||||||
ua_pretty: str = ""
|
|
||||||
#: External https origin of the initial load, "" for direct/none.
|
#: External https origin of the initial load, "" for direct/none.
|
||||||
referer: str = ""
|
referer: str = ""
|
||||||
#: Raw query string of the landing URL (UTM tags can be parsed from it).
|
#: Raw query string of the landing URL (UTM tags can be parsed from it).
|
||||||
query: str = ""
|
query: str = ""
|
||||||
|
|
||||||
|
|
||||||
|
class AbuseHit(msgspec.Struct, omit_defaults=True):
|
||||||
|
"""A request from an IP classified as a scanner/abuser.
|
||||||
|
|
||||||
|
Unlike crawler hits the full request path (query string included) is
|
||||||
|
kept: the interesting part is exactly which paths were probed.
|
||||||
|
``flag`` marks the path that triggered classification; ``is_404``
|
||||||
|
distinguishes 404 responses from document GETs made by the abuser.
|
||||||
|
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
|
||||||
|
"""
|
||||||
|
|
||||||
|
start: datetime
|
||||||
|
#: Full request path including the query string (e.g. "/.env?x=1").
|
||||||
|
path: str
|
||||||
|
#: 6-byte blake3 hash referencing ``Analytics.clients``.
|
||||||
|
client: bytes = b""
|
||||||
|
#: True when this path triggered abuse classification (telltale path
|
||||||
|
#: or the 404 that crossed the threshold).
|
||||||
|
flag: bool = False
|
||||||
|
#: True for 404 responses; false for document GETs from the abuser.
|
||||||
|
is_404: bool = False
|
||||||
|
|
||||||
|
|
||||||
class Analytics(msgspec.Struct, omit_defaults=True):
|
class Analytics(msgspec.Struct, omit_defaults=True):
|
||||||
"""Root of the analytics JSON file. Append-only by design: old data is
|
"""Root of the analytics JSON file. Append-only by design: old data is
|
||||||
dropped by deleting list entries / bucket keys."""
|
dropped by deleting list entries / bucket keys."""
|
||||||
@@ -98,6 +156,12 @@ class Analytics(msgspec.Struct, omit_defaults=True):
|
|||||||
visits: list[Visit] = []
|
visits: list[Visit] = []
|
||||||
#: Document GETs that never produced a ping, treated as crawler/bot hits.
|
#: Document GETs that never produced a ping, treated as crawler/bot hits.
|
||||||
crawlers: list[CrawlerHit] = []
|
crawlers: list[CrawlerHit] = []
|
||||||
|
#: Requests from abusive IPs (see AbuseHit), grouped by IP in the viewer.
|
||||||
|
abuse: list[AbuseHit] = []
|
||||||
|
#: Client metadata keyed by 6-byte blake3 hash.
|
||||||
|
clients: dict[bytes, Client] = {}
|
||||||
|
#: IPs classified as scanners/abusers (keys; values always True).
|
||||||
|
abuse_ips: dict[str, bool] = {}
|
||||||
#: Page transitions per 5-minute bucket (sparse):
|
#: Page transitions per 5-minute bucket (sparse):
|
||||||
#: from -> to -> bucket ISO -> count. ``from`` is the referer origin or
|
#: from -> to -> bucket ISO -> count. ``from`` is the referer origin or
|
||||||
#: "(direct)" for initial loads, a page path for pings.
|
#: "(direct)" for initial loads, a page path for pings.
|
||||||
@@ -124,6 +188,17 @@ def _origin(url: str) -> str | None:
|
|||||||
return f"https://{parsed.netloc}"
|
return f"https://{parsed.netloc}"
|
||||||
|
|
||||||
|
|
||||||
|
def _external_target(url: str) -> str | None:
|
||||||
|
"""A valid https URL (origin or full page), else None."""
|
||||||
|
try:
|
||||||
|
parsed = urlparse(url)
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
|
if parsed.scheme != "https" or not parsed.netloc:
|
||||||
|
return None
|
||||||
|
return url
|
||||||
|
|
||||||
|
|
||||||
_SEGMENT = re.compile(r"[a-z0-9][a-z0-9_-]*")
|
_SEGMENT = re.compile(r"[a-z0-9][a-z0-9_-]*")
|
||||||
|
|
||||||
|
|
||||||
@@ -170,9 +245,62 @@ def _utm_tags(query: str) -> dict[str, str]:
|
|||||||
|
|
||||||
_CRAWLER_TIMEOUT = timedelta(seconds=10)
|
_CRAWLER_TIMEOUT = timedelta(seconds=10)
|
||||||
|
|
||||||
|
#: UAs of JS-running crawlers, which would register as visitors on their
|
||||||
|
#: ping. Anything calling itself a "bot" matches; known crawlers without
|
||||||
|
#: that token (GoogleOther) are listed as extra alternates. No source
|
||||||
|
#: verification: a spoofed bot UA just lands in the crawler list, and
|
||||||
|
#: scanners that probe telltale paths are caught by the abuse rules anyway.
|
||||||
|
_BOT_UA = re.compile(r"bot|googleother", re.IGNORECASE)
|
||||||
|
|
||||||
|
|
||||||
|
def _is_bot_ua(ua: str) -> bool:
|
||||||
|
"""True when the UA claims a crawler identity (Googlebot, Applebot, ...)."""
|
||||||
|
return bool(_BOT_UA.search(ua))
|
||||||
|
|
||||||
|
#: Plain-404 count per IP that classifies it as abuse even without a
|
||||||
|
#: telltale path hit.
|
||||||
|
_ABUSE_404_THRESHOLD = 10
|
||||||
|
|
||||||
|
#: Paths that instantly classify an IP as abuse when they 404: any segment
|
||||||
|
#: starting with a dot ("/.env", "/.git/config") or ending in ".php".
|
||||||
|
_ABUSE_PATH = re.compile(r"(^|/)\.|\.php$", re.IGNORECASE)
|
||||||
|
|
||||||
|
|
||||||
|
def _is_abuse_path(path: str) -> bool:
|
||||||
|
"""Telltale scanner path: dot segment or *.php."""
|
||||||
|
return bool(_ABUSE_PATH.search(path.split("?")[0]))
|
||||||
|
|
||||||
|
|
||||||
|
def _network_ip(ip: str) -> str:
|
||||||
|
"""IPv4 address unchanged, IPv6 collapsed to its /64 network address.
|
||||||
|
|
||||||
|
We hash the network rather than the full address so that clients in the
|
||||||
|
same /64 (a typical end-user allocation) are treated as one visitor.
|
||||||
|
"""
|
||||||
|
if not ip:
|
||||||
|
return ip
|
||||||
|
try:
|
||||||
|
addr = ipaddress.ip_address(ip)
|
||||||
|
except ValueError:
|
||||||
|
return ip
|
||||||
|
if isinstance(addr, ipaddress.IPv6Address):
|
||||||
|
return str(ipaddress.IPv6Network(f"{ip}/64", strict=False).network_address)
|
||||||
|
return ip
|
||||||
|
|
||||||
|
|
||||||
|
def _client_hash(ip: str, ua: str, lang: str) -> bytes:
|
||||||
|
"""6-byte blake3 digest identifying a visitor/client tuple.
|
||||||
|
|
||||||
|
The key is the prettified IP (IPv6 /64), the raw UA string and the
|
||||||
|
extracted language tag, separated by null bytes.
|
||||||
|
"""
|
||||||
|
return blake3.blake3(
|
||||||
|
f"{_network_ip(ip)}\0{ua}\0{lang}".encode()
|
||||||
|
).digest()[:6]
|
||||||
|
|
||||||
|
|
||||||
class Store:
|
class Store:
|
||||||
"""In-memory analytics data plus the (IP, UA) -> visit session map."""
|
"""In-memory analytics data plus the client-hash -> visit session map."""
|
||||||
|
|
||||||
def __init__(self, path: Path) -> None:
|
def __init__(self, path: Path) -> None:
|
||||||
self.path = path
|
self.path = path
|
||||||
@@ -182,8 +310,8 @@ class Store:
|
|||||||
self.data = msgspec.json.decode(path.read_bytes(), type=Analytics)
|
self.data = msgspec.json.decode(path.read_bytes(), type=Analytics)
|
||||||
except msgspec.DecodeError, OSError:
|
except msgspec.DecodeError, OSError:
|
||||||
pass # legacy schema / corrupt or unreadable file: start fresh
|
pass # legacy schema / corrupt or unreadable file: start fresh
|
||||||
#: (ip, user-agent) -> index of the current visit in data.visits
|
#: client hash -> index of the current visit in data.visits
|
||||||
self.sessions: dict[tuple[str, str], int] = {}
|
self.sessions: dict[bytes, int] = {}
|
||||||
#: ip -> external https origin of the latest document GET carrying
|
#: ip -> external https origin of the latest document GET carrying
|
||||||
#: one, stashed for the visit the client's initial ping starts.
|
#: one, stashed for the visit the client's initial ping starts.
|
||||||
#: Internal or absent referers never touch the table.
|
#: Internal or absent referers never touch the table.
|
||||||
@@ -196,6 +324,26 @@ class Store:
|
|||||||
#: Document GETs that have not yet been matched by a ping. Kept
|
#: Document GETs that have not yet been matched by a ping. Kept
|
||||||
#: in RAM only; expired entries are written to ``data.crawlers``.
|
#: in RAM only; expired entries are written to ``data.crawlers``.
|
||||||
self.pending_crawlers: list[CrawlerHit] = []
|
self.pending_crawlers: list[CrawlerHit] = []
|
||||||
|
#: ip -> number of plain (non-telltale) 404s seen, in RAM only;
|
||||||
|
#: reaching ``_ABUSE_404_THRESHOLD`` classifies the IP as abuse.
|
||||||
|
self.not_found_counts: dict[str, int] = {}
|
||||||
|
#: Callables to notify when persisted data changes. Registered by the
|
||||||
|
#: analytics WebSocket broadcaster.
|
||||||
|
self._on_change: list[Callable[[], None]] = []
|
||||||
|
|
||||||
|
def subscribe(self, callback: Callable[[], None]) -> None:
|
||||||
|
"""Register a callback to be called after every persisted change."""
|
||||||
|
if callback not in self._on_change:
|
||||||
|
self._on_change.append(callback)
|
||||||
|
|
||||||
|
def unsubscribe(self, callback: Callable[[], None]) -> None:
|
||||||
|
"""Remove a previously registered change callback."""
|
||||||
|
with suppress(ValueError):
|
||||||
|
self._on_change.remove(callback)
|
||||||
|
|
||||||
|
def _notify(self) -> None:
|
||||||
|
for callback in self._on_change:
|
||||||
|
callback()
|
||||||
|
|
||||||
def _save(self) -> None:
|
def _save(self) -> None:
|
||||||
"""Rewrite the JSON file atomically (temp file + rename)."""
|
"""Rewrite the JSON file atomically (temp file + rename)."""
|
||||||
@@ -208,21 +356,29 @@ class Store:
|
|||||||
os.replace(tmp, self.path)
|
os.replace(tmp, self.path)
|
||||||
except OSError:
|
except OSError:
|
||||||
pass # analytics must never break page serving
|
pass # analytics must never break page serving
|
||||||
|
else:
|
||||||
|
self._notify()
|
||||||
|
|
||||||
def _flush_crawlers(self, now: datetime | None = None) -> None:
|
def _flush_crawlers(self, now: datetime | None = None) -> list[bytes]:
|
||||||
"""Move expired pending crawler hits into persistent ``data.crawlers``."""
|
"""Move expired pending crawler hits into persistent ``data.crawlers``.
|
||||||
|
|
||||||
|
Returns the client hashes of the newly flushed hits so callers can
|
||||||
|
schedule async enrichment.
|
||||||
|
"""
|
||||||
if not self.pending_crawlers:
|
if not self.pending_crawlers:
|
||||||
return
|
return []
|
||||||
now = now or datetime.now(UTC)
|
now = now or datetime.now(UTC)
|
||||||
cutoff = now - _CRAWLER_TIMEOUT
|
cutoff = now - _CRAWLER_TIMEOUT
|
||||||
expired: list[CrawlerHit] = []
|
expired: list[CrawlerHit] = []
|
||||||
remaining: list[CrawlerHit] = []
|
remaining: list[CrawlerHit] = []
|
||||||
for hit in self.pending_crawlers:
|
for hit in self.pending_crawlers:
|
||||||
(expired if hit.start <= cutoff else remaining).append(hit)
|
(expired if hit.start <= cutoff else remaining).append(hit)
|
||||||
if expired:
|
if not expired:
|
||||||
self.pending_crawlers = remaining
|
return []
|
||||||
self.data.crawlers.extend(expired)
|
self.pending_crawlers = remaining
|
||||||
self._save()
|
self.data.crawlers.extend(expired)
|
||||||
|
self._save()
|
||||||
|
return [hit.client for hit in expired]
|
||||||
|
|
||||||
def _count(self, table: dict[str, int], key: str) -> None:
|
def _count(self, table: dict[str, int], key: str) -> None:
|
||||||
table[key] = table.get(key, 0) + 1
|
table[key] = table.get(key, 0) + 1
|
||||||
@@ -232,15 +388,188 @@ class Store:
|
|||||||
buckets = self.data.transitions.setdefault(fr, {}).setdefault(to, {})
|
buckets = self.data.transitions.setdefault(fr, {}).setdefault(to, {})
|
||||||
self._count(buckets, _bucket(now))
|
self._count(buckets, _bucket(now))
|
||||||
|
|
||||||
|
def _uncount(self, table: dict[str, int], key: str) -> None:
|
||||||
|
"""Reverse one ``_count``: decrement and drop empty keys."""
|
||||||
|
if key in table:
|
||||||
|
table[key] -= 1
|
||||||
|
if table[key] <= 0:
|
||||||
|
del table[key]
|
||||||
|
|
||||||
|
def _remove_visit(self, index: int) -> None:
|
||||||
|
"""Delete a visit and reverse the counts its creation recorded.
|
||||||
|
|
||||||
|
Used when a known visitor turns out to be an admin (hide=1 ping):
|
||||||
|
the session is scrubbed from the stats. Views/transitions logged
|
||||||
|
by later pings inside the visit lack per-event timestamps and are
|
||||||
|
left as-is.
|
||||||
|
"""
|
||||||
|
visit = self.data.visits[index]
|
||||||
|
bucket = _bucket(visit.start)
|
||||||
|
self._uncount(self.data.site_visits, bucket)
|
||||||
|
views = self.data.views.get(visit.entry)
|
||||||
|
if views is not None:
|
||||||
|
self._uncount(views, bucket)
|
||||||
|
if not views:
|
||||||
|
del self.data.views[visit.entry]
|
||||||
|
fr_map = self.data.transitions.get(visit.referer or "(direct)")
|
||||||
|
if fr_map is not None:
|
||||||
|
buckets = fr_map.get(visit.entry)
|
||||||
|
if buckets is not None:
|
||||||
|
self._uncount(buckets, bucket)
|
||||||
|
if not buckets:
|
||||||
|
del fr_map[visit.entry]
|
||||||
|
if not fr_map:
|
||||||
|
del self.data.transitions[visit.referer or "(direct)"]
|
||||||
|
del self.data.visits[index]
|
||||||
|
# Sessions store list indices; shift the ones past the removed visit.
|
||||||
|
for key, i in list(self.sessions.items()):
|
||||||
|
if i > index:
|
||||||
|
self.sessions[key] = i - 1
|
||||||
|
|
||||||
|
def _client_ip(self, client_hash: bytes) -> str:
|
||||||
|
"""Return the IP stored for ``client_hash``, or "" if missing."""
|
||||||
|
client = self.data.clients.get(client_hash)
|
||||||
|
return client.ip if client else ""
|
||||||
|
|
||||||
|
def _ensure_client(
|
||||||
|
self,
|
||||||
|
ip: str,
|
||||||
|
ua: str,
|
||||||
|
lang: str,
|
||||||
|
*,
|
||||||
|
country: str = "",
|
||||||
|
) -> bytes:
|
||||||
|
"""Get or create a ``Client`` record; return its 6-byte hash."""
|
||||||
|
h = _client_hash(ip, ua, lang)
|
||||||
|
if h not in self.data.clients:
|
||||||
|
self.data.clients[h] = Client(
|
||||||
|
ip=ip,
|
||||||
|
ua=ua,
|
||||||
|
ua_pretty=_compact_user_agent(ua),
|
||||||
|
lang=lang,
|
||||||
|
country=country,
|
||||||
|
)
|
||||||
|
self._save()
|
||||||
|
return h
|
||||||
|
|
||||||
|
def enrich_client(
|
||||||
|
self,
|
||||||
|
client_hash: bytes,
|
||||||
|
*,
|
||||||
|
host: str = "",
|
||||||
|
country: str = "",
|
||||||
|
city: str = "",
|
||||||
|
) -> None:
|
||||||
|
"""Fill in host/geoip fields on a client record after async lookups."""
|
||||||
|
client = self.data.clients.get(client_hash)
|
||||||
|
if client is None:
|
||||||
|
return
|
||||||
|
changed = False
|
||||||
|
if host and not client.host:
|
||||||
|
client.host = host
|
||||||
|
changed = True
|
||||||
|
if country:
|
||||||
|
client.country = country
|
||||||
|
changed = True
|
||||||
|
if city:
|
||||||
|
client.city = city
|
||||||
|
changed = True
|
||||||
|
if changed:
|
||||||
|
self._save()
|
||||||
|
|
||||||
|
def _abuse_hit(
|
||||||
|
self,
|
||||||
|
client_hash: bytes,
|
||||||
|
path: str,
|
||||||
|
start: datetime | None = None,
|
||||||
|
*,
|
||||||
|
flag: bool = False,
|
||||||
|
is_404: bool = False,
|
||||||
|
) -> None:
|
||||||
|
"""Append one abuse hit referencing a client by hash."""
|
||||||
|
self.data.abuse.append(
|
||||||
|
AbuseHit(
|
||||||
|
start=start or datetime.now(UTC),
|
||||||
|
path=path,
|
||||||
|
client=client_hash,
|
||||||
|
flag=flag,
|
||||||
|
is_404=is_404,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
def classify_abuse(
|
||||||
|
self,
|
||||||
|
ip: str,
|
||||||
|
client_hash: bytes,
|
||||||
|
path: str,
|
||||||
|
*,
|
||||||
|
flag: bool = False,
|
||||||
|
is_404: bool = False,
|
||||||
|
) -> None:
|
||||||
|
"""Classify an IP as a scanner/abuser and record the triggering hit.
|
||||||
|
|
||||||
|
All earlier crawler hits from the same IP (persisted and pending)
|
||||||
|
are moved to the abuse list — a random-UA scanner must not pollute
|
||||||
|
the crawler stats of the legitimate bots it impersonates.
|
||||||
|
"""
|
||||||
|
if ip not in self.data.abuse_ips:
|
||||||
|
self.data.abuse_ips[ip] = True
|
||||||
|
moved = [h for h in self.data.crawlers if self._client_ip(h.client) == ip]
|
||||||
|
if moved:
|
||||||
|
self.data.crawlers = [h for h in self.data.crawlers if self._client_ip(h.client) != ip]
|
||||||
|
for h in moved:
|
||||||
|
self._abuse_hit(
|
||||||
|
h.client,
|
||||||
|
h.entry + (f"?{h.query}" if h.query else ""),
|
||||||
|
start=h.start,
|
||||||
|
)
|
||||||
|
pending = [h for h in self.pending_crawlers if self._client_ip(h.client) == ip]
|
||||||
|
if pending:
|
||||||
|
self.pending_crawlers = [h for h in self.pending_crawlers if self._client_ip(h.client) != ip]
|
||||||
|
for h in pending:
|
||||||
|
self._abuse_hit(
|
||||||
|
h.client,
|
||||||
|
h.entry + (f"?{h.query}" if h.query else ""),
|
||||||
|
start=h.start,
|
||||||
|
)
|
||||||
|
self._abuse_hit(client_hash, path, flag=flag, is_404=is_404)
|
||||||
|
self._save()
|
||||||
|
|
||||||
|
def track_404(
|
||||||
|
self,
|
||||||
|
ip: str,
|
||||||
|
ua: str,
|
||||||
|
path: str,
|
||||||
|
accept_language: str = "",
|
||||||
|
) -> bytes:
|
||||||
|
"""Record a 404 response for ``path`` (full path, query included).
|
||||||
|
|
||||||
|
A telltale path (dot segment or *.php) classifies the IP as abuse
|
||||||
|
immediately; enough plain 404s from one IP do too. Hits from
|
||||||
|
already-classified IPs go straight to the abuse list.
|
||||||
|
|
||||||
|
Returns the client hash so callers can schedule async enrichment.
|
||||||
|
"""
|
||||||
|
lang, country = _parse_accept_language(accept_language)
|
||||||
|
client_hash = self._ensure_client(ip, ua, lang, country=country)
|
||||||
|
if ip in self.data.abuse_ips:
|
||||||
|
self._abuse_hit(client_hash, path, flag=_is_abuse_path(path), is_404=True)
|
||||||
|
self._save()
|
||||||
|
return client_hash
|
||||||
|
if _is_abuse_path(path):
|
||||||
|
self.classify_abuse(ip, client_hash, path, flag=True, is_404=True)
|
||||||
|
return client_hash
|
||||||
|
self.not_found_counts[ip] = self.not_found_counts.get(ip, 0) + 1
|
||||||
|
if self.not_found_counts[ip] >= _ABUSE_404_THRESHOLD:
|
||||||
|
self.classify_abuse(ip, client_hash, path, flag=True, is_404=True)
|
||||||
|
return client_hash
|
||||||
|
return client_hash
|
||||||
|
|
||||||
def _new_visit(
|
def _new_visit(
|
||||||
self,
|
self,
|
||||||
entry: str,
|
entry: str,
|
||||||
referer: str,
|
referer: str,
|
||||||
key: tuple[str, str],
|
client_hash: bytes,
|
||||||
ip: str = "",
|
|
||||||
lang: str = "",
|
|
||||||
country: str = "",
|
|
||||||
ua: str = "",
|
|
||||||
utm: dict[str, str] | None = None,
|
utm: dict[str, str] | None = None,
|
||||||
) -> Visit:
|
) -> Visit:
|
||||||
now = datetime.now(UTC)
|
now = datetime.now(UTC)
|
||||||
@@ -248,50 +577,25 @@ class Store:
|
|||||||
start=now,
|
start=now,
|
||||||
entry=entry,
|
entry=entry,
|
||||||
referer=referer,
|
referer=referer,
|
||||||
ip=ip,
|
client=client_hash,
|
||||||
lang=lang,
|
|
||||||
country=country,
|
|
||||||
ua=ua,
|
|
||||||
ua_pretty=_compact_user_agent(ua),
|
|
||||||
utm=utm or {},
|
utm=utm or {},
|
||||||
)
|
)
|
||||||
self.data.visits.append(visit)
|
self.data.visits.append(visit)
|
||||||
self.sessions[key] = len(self.data.visits) - 1
|
self.sessions[client_hash] = len(self.data.visits) - 1
|
||||||
self._count(self.data.site_visits, _bucket(now))
|
self._count(self.data.site_visits, _bucket(now))
|
||||||
self._count(self.data.views.setdefault(entry, {}), _bucket(now))
|
self._count(self.data.views.setdefault(entry, {}), _bucket(now))
|
||||||
self._count_transition(referer or "(direct)", entry, now)
|
self._count_transition(referer or "(direct)", entry, now)
|
||||||
return visit
|
return visit
|
||||||
|
|
||||||
def enrich_visit(
|
|
||||||
self,
|
|
||||||
index: int,
|
|
||||||
*,
|
|
||||||
host: str = "",
|
|
||||||
country: str = "",
|
|
||||||
) -> None:
|
|
||||||
"""Fill in host/geoip fields on an existing visit after async lookups."""
|
|
||||||
if index < 0 or index >= len(self.data.visits):
|
|
||||||
return
|
|
||||||
visit = self.data.visits[index]
|
|
||||||
changed = False
|
|
||||||
if host and not visit.host:
|
|
||||||
visit.host = host
|
|
||||||
changed = True
|
|
||||||
if country:
|
|
||||||
visit.country = country
|
|
||||||
changed = True
|
|
||||||
if changed:
|
|
||||||
self._save()
|
|
||||||
|
|
||||||
def track_entry(
|
def track_entry(
|
||||||
self,
|
self,
|
||||||
referer: str,
|
referer: str,
|
||||||
own_origin: str,
|
own_origin: str,
|
||||||
ip: str,
|
ip: str,
|
||||||
ua: str,
|
ua: str,
|
||||||
entry: str,
|
full_path: str,
|
||||||
query: str = "",
|
accept_language: str = "",
|
||||||
) -> None:
|
) -> list[bytes]:
|
||||||
"""Stash the entry referer/UTM tags and queue a pending crawler hit.
|
"""Stash the entry referer/UTM tags and queue a pending crawler hit.
|
||||||
|
|
||||||
Nothing is counted here — the client's initial /_a ping starts the
|
Nothing is counted here — the client's initial /_a ping starts the
|
||||||
@@ -302,11 +606,28 @@ class Store:
|
|||||||
does not erase an earlier tagged landing.
|
does not erase an earlier tagged landing.
|
||||||
|
|
||||||
Every document GET is also queued as a pending crawler hit. If a ping
|
Every document GET is also queued as a pending crawler hit. If a ping
|
||||||
from the same (IP, UA) pair arrives within ``_CRAWLER_TIMEOUT``, the
|
from the same client arrives within ``_CRAWLER_TIMEOUT``, the hit is
|
||||||
hit is discarded; otherwise it is flushed to ``data.crawlers``.
|
discarded; otherwise it is flushed to ``data.crawlers``. The
|
||||||
|
Accept-Language header is stored on the client record immediately;
|
||||||
|
host/geoip are filled in later by async enrichment.
|
||||||
|
|
||||||
|
GETs from IPs already classified as abuse are recorded as abuse hits
|
||||||
|
with the full request path (query string included).
|
||||||
|
|
||||||
|
Returns the client hashes of any hits flushed to persistent storage,
|
||||||
|
so callers can schedule async enrichment.
|
||||||
"""
|
"""
|
||||||
|
entry = full_path.split("?")[0]
|
||||||
|
query = full_path.split("?", 1)[1] if "?" in full_path else ""
|
||||||
|
lang, country = _parse_accept_language(accept_language)
|
||||||
|
client_hash = self._ensure_client(ip, ua, lang, country=country)
|
||||||
|
if ip in self.data.abuse_ips:
|
||||||
|
flushed = self._flush_crawlers()
|
||||||
|
self._abuse_hit(client_hash, full_path, is_404=False, flag=False)
|
||||||
|
self._save()
|
||||||
|
return flushed
|
||||||
now = datetime.now(UTC)
|
now = datetime.now(UTC)
|
||||||
self._flush_crawlers(now)
|
flushed = self._flush_crawlers(now)
|
||||||
if referer:
|
if referer:
|
||||||
origin = _origin(referer)
|
origin = _origin(referer)
|
||||||
if origin is not None and origin != own_origin:
|
if origin is not None and origin != own_origin:
|
||||||
@@ -318,61 +639,106 @@ class Store:
|
|||||||
CrawlerHit(
|
CrawlerHit(
|
||||||
start=now,
|
start=now,
|
||||||
entry=entry,
|
entry=entry,
|
||||||
ip=ip,
|
client=client_hash,
|
||||||
ua=ua,
|
|
||||||
ua_pretty=_compact_user_agent(ua),
|
|
||||||
referer=self.pending_referers.get(ip, ""),
|
referer=self.pending_referers.get(ip, ""),
|
||||||
query=query,
|
query=query,
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
return flushed
|
||||||
|
|
||||||
|
def _add_read(self, client_hash: bytes, path: str, seconds: int) -> None:
|
||||||
|
"""Add ``seconds`` of reading time for ``path`` to the current visit."""
|
||||||
|
if seconds <= 0:
|
||||||
|
return
|
||||||
|
index = self.sessions.get(client_hash)
|
||||||
|
if index is None or index >= len(self.data.visits):
|
||||||
|
return
|
||||||
|
visit = self.data.visits[index]
|
||||||
|
visit.read[path] = visit.read.get(path, 0) + seconds
|
||||||
|
|
||||||
def ping(
|
def ping(
|
||||||
self,
|
self,
|
||||||
from_: str,
|
from_: str,
|
||||||
to: str,
|
to: str | None,
|
||||||
ip: str,
|
ip: str,
|
||||||
ua: str,
|
ua: str,
|
||||||
accept_language: str = "",
|
accept_language: str = "",
|
||||||
) -> int | None:
|
hide: bool = False,
|
||||||
"""Record a client navigation ping ({from, to} from pagerite.js).
|
read: int = 0,
|
||||||
|
) -> tuple[int | None, list[bytes]]:
|
||||||
|
"""Record a client navigation ping ({from, to, read} from pagerite.js).
|
||||||
|
|
||||||
|
``to`` is an internal path ("/...") or an https URL for exit links; a
|
||||||
|
missing/empty ``to`` means the page is being closed and only the
|
||||||
|
``read`` time should be recorded. The transition is always counted when
|
||||||
|
``to`` is present; the trail only grows on first sight of a page within
|
||||||
|
the visit. ``read`` is the active time (seconds) spent on ``from_``.
|
||||||
|
|
||||||
``to`` is an internal path ("/...") or an https origin for exit
|
|
||||||
links; anything else is ignored. The transition is always counted;
|
|
||||||
the trail only grows on first sight of a page within the visit.
|
|
||||||
A ping with no known session starts a fresh visit, consuming the
|
A ping with no known session starts a fresh visit, consuming the
|
||||||
referer and UTM tags stashed by the document GET if there are any.
|
referer and UTM tags stashed by the document GET if there are any.
|
||||||
|
|
||||||
Returns the index of the new visit when one is created, so callers
|
``hide`` is set by admin clients: the ping cancels pending crawler
|
||||||
can enrich it later with non-blocking lookups (host, geoip country).
|
hits as usual, and any existing visit for this client session is
|
||||||
|
removed from the stats (the admin browsed anonymously before logging
|
||||||
|
in). Nothing new is recorded.
|
||||||
|
|
||||||
|
Pings from IPs classified as abuse, and pings whose User-Agent
|
||||||
|
claims a JS-running crawler identity (``_is_bot_ua``), are ignored
|
||||||
|
entirely — the crawler's pending hits stay queued and flush to
|
||||||
|
``data.crawlers`` normally.
|
||||||
|
|
||||||
|
Returns the index of the new visit when one is created (or None) and
|
||||||
|
the client hashes of any crawler hits flushed by this call, so callers
|
||||||
|
can schedule async enrichment (host, geoip country/city).
|
||||||
"""
|
"""
|
||||||
self._flush_crawlers()
|
flushed = self._flush_crawlers()
|
||||||
# A real visitor ping cancels any pending crawler hits from this
|
lang, country = _parse_accept_language(accept_language)
|
||||||
# (IP, UA) pair.
|
client_hash = _client_hash(ip, ua, lang)
|
||||||
|
if hide:
|
||||||
|
# Admin ping: cancel pending crawler hits and scrub the session.
|
||||||
|
self.pending_crawlers = [
|
||||||
|
hit for hit in self.pending_crawlers if hit.client == client_hash
|
||||||
|
]
|
||||||
|
index = self.sessions.pop(client_hash, None)
|
||||||
|
if index is not None and index < len(self.data.visits):
|
||||||
|
self._remove_visit(index)
|
||||||
|
self._save()
|
||||||
|
return None, flushed
|
||||||
|
if ip in self.data.abuse_ips:
|
||||||
|
return None, flushed
|
||||||
|
if _is_bot_ua(ua):
|
||||||
|
# A JS-running crawler (Googlebot, GoogleOther, Applebot execute
|
||||||
|
# JS and ping): never a visit. Its pending crawler hits are
|
||||||
|
# kept and flush to ``data.crawlers`` normally.
|
||||||
|
return None, flushed
|
||||||
|
# A real visitor ping cancels any pending crawler hits from this client.
|
||||||
self.pending_crawlers = [
|
self.pending_crawlers = [
|
||||||
hit for hit in self.pending_crawlers if not (hit.ip == ip and hit.ua == ua)
|
hit for hit in self.pending_crawlers if hit.client != client_hash
|
||||||
]
|
]
|
||||||
|
fr_path = _internal_path(from_) if from_ else ""
|
||||||
|
if fr_path and read > 0:
|
||||||
|
self._add_read(client_hash, fr_path, read)
|
||||||
|
if not to:
|
||||||
|
if read > 0:
|
||||||
|
self._save()
|
||||||
|
return None, flushed
|
||||||
if to.startswith("/") and not to.startswith("//"):
|
if to.startswith("/") and not to.startswith("//"):
|
||||||
target = _internal_path(to) or ""
|
target = _internal_path(to) or ""
|
||||||
else:
|
else:
|
||||||
target = _origin(to) or ""
|
target = _external_target(to) or ""
|
||||||
if not target or (not to.startswith("/") and target != to):
|
if not target:
|
||||||
return None
|
return None, flushed
|
||||||
key = (ip, ua)
|
index = self.sessions.get(client_hash)
|
||||||
index = self.sessions.get(key)
|
fr = fr_path or "(direct)"
|
||||||
fr = (_internal_path(from_) or "(direct)") if from_ else "(direct)"
|
|
||||||
if index is None or index >= len(self.data.visits):
|
if index is None or index >= len(self.data.visits):
|
||||||
# No known session: the initial ping of a fresh page load (or
|
# No known session: the initial ping of a fresh page load (or
|
||||||
# missing data after a server restart) — start a visit.
|
# missing data after a server restart) — start a visit.
|
||||||
lang, country = _parse_accept_language(accept_language)
|
|
||||||
index = len(self.data.visits)
|
index = len(self.data.visits)
|
||||||
|
self._ensure_client(ip, ua, lang, country=country)
|
||||||
self._new_visit(
|
self._new_visit(
|
||||||
target,
|
target,
|
||||||
self.pending_referers.pop(ip, ""),
|
self.pending_referers.pop(ip, ""),
|
||||||
key,
|
client_hash,
|
||||||
ip=ip,
|
|
||||||
lang=lang,
|
|
||||||
country=country,
|
|
||||||
ua=ua,
|
|
||||||
utm=self.pending_utms.pop(ip, {}),
|
utm=self.pending_utms.pop(ip, {}),
|
||||||
)
|
)
|
||||||
else:
|
else:
|
||||||
@@ -385,4 +751,5 @@ class Store:
|
|||||||
if visit.entry != target and target not in visit.trail:
|
if visit.entry != target and target not in visit.trail:
|
||||||
visit.trail.append(target)
|
visit.trail.append(target)
|
||||||
self._save()
|
self._save()
|
||||||
return index if index is not None and index < len(self.data.visits) else None
|
visit_index = index if index is not None and index < len(self.data.visits) else None
|
||||||
|
return visit_index, flushed
|
||||||
|
|||||||
+344
-38
@@ -27,14 +27,16 @@ from email.utils import format_datetime
|
|||||||
from functools import lru_cache
|
from functools import lru_cache
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from urllib.parse import urlparse
|
from urllib.parse import urlparse
|
||||||
|
from xml.sax.saxutils import escape as xml_escape
|
||||||
|
|
||||||
import blake3
|
import blake3
|
||||||
import msgspec
|
import msgspec
|
||||||
from fastapi import FastAPI, HTTPException, Request, WebSocket, WebSocketDisconnect
|
from fastapi import FastAPI, HTTPException, Request, WebSocket, WebSocketDisconnect
|
||||||
from fastapi.responses import HTMLResponse, RedirectResponse, Response
|
from fastapi.responses import RedirectResponse, Response
|
||||||
from fastapi_vue import Frontend
|
from fastapi_vue import Frontend
|
||||||
from kanta import Kanta
|
from kanta import Kanta
|
||||||
from pydantic import BaseModel
|
from pydantic import BaseModel
|
||||||
|
from zstandard import ZstdCompressor
|
||||||
|
|
||||||
from pagerite import analytics, seed, views
|
from pagerite import analytics, seed, views
|
||||||
from pagerite.__main__ import DEVMODE
|
from pagerite.__main__ import DEVMODE
|
||||||
@@ -57,6 +59,10 @@ ANALYTICS_PATH = Path(
|
|||||||
)
|
)
|
||||||
analytics_store = analytics.Store(ANALYTICS_PATH)
|
analytics_store = analytics.Store(ANALYTICS_PATH)
|
||||||
|
|
||||||
|
# Live WebSocket clients for the analytics stream.
|
||||||
|
_analytics_ws_clients: set[WebSocket] = set()
|
||||||
|
_analytics_broadcast_task: asyncio.Task | None = None
|
||||||
|
|
||||||
|
|
||||||
# Repository root from this file's location (pagerite/app.py -> ..).
|
# Repository root from this file's location (pagerite/app.py -> ..).
|
||||||
_REPO_ROOT = Path(__file__).resolve().parent.parent
|
_REPO_ROOT = Path(__file__).resolve().parent.parent
|
||||||
@@ -121,6 +127,26 @@ class GeoIP:
|
|||||||
pass
|
pass
|
||||||
return ""
|
return ""
|
||||||
|
|
||||||
|
def city(self, ip: str) -> str:
|
||||||
|
"""City name for ``ip``, or "" when unavailable.
|
||||||
|
|
||||||
|
GeoIP sometimes appends district names in parentheses (e.g.
|
||||||
|
"Berlin (Bezirk Tempelhof-Schöneberg)"); those are stripped before
|
||||||
|
the value is stored.
|
||||||
|
"""
|
||||||
|
if not ip or self._reader is None:
|
||||||
|
return ""
|
||||||
|
try:
|
||||||
|
rec = self._reader.get(ip)
|
||||||
|
if rec:
|
||||||
|
city = (rec.get("city") or {}).get("names", {}).get("en", "")
|
||||||
|
if city:
|
||||||
|
city = re.sub(r"\s*\([^)]*\)", "", city).strip()
|
||||||
|
return city
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
_geoip = GeoIP()
|
_geoip = GeoIP()
|
||||||
|
|
||||||
@@ -231,7 +257,9 @@ async def lifespan(_app: FastAPI) -> AsyncIterator[None]:
|
|||||||
# Decompress/open the DB-IP MMDB once at startup. Lookups are then
|
# Decompress/open the DB-IP MMDB once at startup. Lookups are then
|
||||||
# read-only and safe to run in background ``to_thread`` workers.
|
# read-only and safe to run in background ``to_thread`` workers.
|
||||||
await asyncio.to_thread(_geoip._load)
|
await asyncio.to_thread(_geoip._load)
|
||||||
|
analytics_store.subscribe(_schedule_analytics_broadcast)
|
||||||
yield
|
yield
|
||||||
|
analytics_store.unsubscribe(_schedule_analytics_broadcast)
|
||||||
await kanta.close()
|
await kanta.close()
|
||||||
|
|
||||||
|
|
||||||
@@ -255,6 +283,79 @@ async def _headers(request: Request, call_next) -> Response:
|
|||||||
return response
|
return response
|
||||||
|
|
||||||
|
|
||||||
|
# Dynamic HTML is compressed per request at level 9 (static assets are
|
||||||
|
# already pre-compressed by fastapi-vue's Frontend).
|
||||||
|
_zstd = ZstdCompressor(9)
|
||||||
|
|
||||||
|
|
||||||
|
def _render_html(kind: str, path: str, base_url: str) -> str:
|
||||||
|
"""Render one of the generated pages (see _html_response)."""
|
||||||
|
if kind == "page":
|
||||||
|
return views.render_page(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html, base_url)
|
||||||
|
if kind == "category":
|
||||||
|
return views.render_category(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html)
|
||||||
|
if kind == "not-found":
|
||||||
|
return views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html)
|
||||||
|
return views.render_analytics(data.menu, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html)
|
||||||
|
|
||||||
|
|
||||||
|
@lru_cache(maxsize=128)
|
||||||
|
def _cached_body(kind: str, path: str, base_url: str, version: int, zstd: bool) -> bytes:
|
||||||
|
"""Rendered page body. Every input the output depends on is in the key:
|
||||||
|
data.version bumps on any content/settings change, base_url feeds the
|
||||||
|
social meta URLs, and zstd selects the stored encoding (both variants
|
||||||
|
are cached rather than re-compressed).
|
||||||
|
"""
|
||||||
|
body = _render_html(kind, path, base_url).encode()
|
||||||
|
return _zstd.compress(body) if zstd else body
|
||||||
|
|
||||||
|
|
||||||
|
def _html_response(
|
||||||
|
request: Request,
|
||||||
|
kind: str,
|
||||||
|
path: str,
|
||||||
|
status_code: int = 200,
|
||||||
|
headers: dict | None = None,
|
||||||
|
etag: bool = False,
|
||||||
|
) -> Response:
|
||||||
|
"""Response for a generated page, zstd-compressed when the client
|
||||||
|
accepts it (no gzip fallback).
|
||||||
|
|
||||||
|
Done per handler rather than in middleware so that Frontend's
|
||||||
|
already-compressed asset responses are never touched. The ETag stays
|
||||||
|
identical across encodings (revalidation compares it before
|
||||||
|
compression); ``vary: accept-encoding`` keeps caches from mixing the
|
||||||
|
representations. In dev the cache is bypassed so theme/design edits on
|
||||||
|
disk apply immediately.
|
||||||
|
|
||||||
|
``etag=True`` derives the validator from a blake3 hash of the
|
||||||
|
(uncompressed) body — for pages like /_a that have no Node whose
|
||||||
|
modified timestamp could serve as one — and answers matching
|
||||||
|
if-none-match revalidations with a 304.
|
||||||
|
"""
|
||||||
|
zstd = "zstd" in request.headers.get("accept-encoding", "")
|
||||||
|
# Absolute social/canonical URLs use the learned public origin; until
|
||||||
|
# an admin visit teaches it, fall back to the request's own base URL.
|
||||||
|
base_url = data.site_url or str(request.base_url).rstrip("/")
|
||||||
|
if DEVMODE:
|
||||||
|
identity = _render_html(kind, path, base_url).encode()
|
||||||
|
body = _zstd.compress(identity) if zstd else identity
|
||||||
|
else:
|
||||||
|
identity = _cached_body(kind, path, base_url, data.version, False)
|
||||||
|
body = _cached_body(kind, path, base_url, data.version, True) if zstd else identity
|
||||||
|
h = dict(headers or {})
|
||||||
|
if zstd:
|
||||||
|
h["vary"] = "accept-encoding"
|
||||||
|
if etag:
|
||||||
|
tag = f'"{blake3.blake3(identity).hexdigest()[:32]}"'
|
||||||
|
h["etag"] = tag
|
||||||
|
if request.headers.get("if-none-match") == tag:
|
||||||
|
return Response(status_code=304, headers=h)
|
||||||
|
if zstd:
|
||||||
|
h["content-encoding"] = "zstd"
|
||||||
|
return Response(body, status_code, h, media_type="text/html")
|
||||||
|
|
||||||
|
|
||||||
class PageIn(BaseModel):
|
class PageIn(BaseModel):
|
||||||
"""Payload for creating or replacing a page."""
|
"""Payload for creating or replacing a page."""
|
||||||
|
|
||||||
@@ -408,6 +509,33 @@ async def put_settings(settings: SettingsIn) -> None:
|
|||||||
data.version += 1
|
data.version += 1
|
||||||
|
|
||||||
|
|
||||||
|
class SiteUrlIn(BaseModel):
|
||||||
|
"""Payload for learning the site's public origin."""
|
||||||
|
|
||||||
|
url: str
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/_api/site-url", status_code=204)
|
||||||
|
async def learn_site_url(payload: SiteUrlIn) -> None:
|
||||||
|
"""Learn the site's public origin (scheme + host) from an admin browser.
|
||||||
|
|
||||||
|
pagerite.js reports location.origin once an admin session is detected:
|
||||||
|
unlike request Host headers it reflects the real public scheme and host
|
||||||
|
even behind reverse proxies, with zero manual configuration. Stored in
|
||||||
|
the database with a version bump so cached pages re-render with correct
|
||||||
|
absolute social/canonical URLs.
|
||||||
|
"""
|
||||||
|
url = payload.url.rstrip("/")
|
||||||
|
parsed = urlparse(url)
|
||||||
|
if parsed.scheme not in ("http", "https") or not parsed.netloc or parsed.path:
|
||||||
|
raise HTTPException(400, "not an origin")
|
||||||
|
if url == data.site_url:
|
||||||
|
return
|
||||||
|
with kanta.transaction("learn site url"):
|
||||||
|
data.site_url = url
|
||||||
|
data.version += 1
|
||||||
|
|
||||||
|
|
||||||
@app.put("/_api/settings/favicon")
|
@app.put("/_api/settings/favicon")
|
||||||
async def put_favicon(request: Request) -> dict[str, str]:
|
async def put_favicon(request: Request) -> dict[str, str]:
|
||||||
"""Upload a favicon into the content-addressed store and activate it.
|
"""Upload a favicon into the content-addressed store and activate it.
|
||||||
@@ -587,6 +715,12 @@ def _client_ip(request: Request) -> str:
|
|||||||
return forwarded or (request.client.host if request.client else "")
|
return forwarded or (request.client.host if request.client else "")
|
||||||
|
|
||||||
|
|
||||||
|
def _query_suffix(request: Request) -> str:
|
||||||
|
"""The request's query string as a "?..." suffix, or "" when absent."""
|
||||||
|
query = str(request.url.query)
|
||||||
|
return f"?{query}" if query else ""
|
||||||
|
|
||||||
|
|
||||||
@lru_cache(maxsize=4096)
|
@lru_cache(maxsize=4096)
|
||||||
def _cached_ptr(ip: str) -> str:
|
def _cached_ptr(ip: str) -> str:
|
||||||
"""Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or ""."""
|
"""Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or ""."""
|
||||||
@@ -615,35 +749,84 @@ async def _geoip_country(ip: str) -> str:
|
|||||||
return await asyncio.to_thread(_geoip.country, ip)
|
return await asyncio.to_thread(_geoip.country, ip)
|
||||||
|
|
||||||
|
|
||||||
async def _enrich_visit(index: int, ip: str) -> None:
|
async def _geoip_city(ip: str) -> str:
|
||||||
"""Run non-blocking reverse-DNS and geoip enrichment for a new visit."""
|
"""Async wrapper around the DB-IP MMDB city lookup."""
|
||||||
if not ip:
|
return await asyncio.to_thread(_geoip.city, ip)
|
||||||
|
|
||||||
|
|
||||||
|
async def _enrich_client(client_hash: bytes) -> None:
|
||||||
|
"""Run non-blocking reverse-DNS and geoip enrichment for a client."""
|
||||||
|
client = analytics_store.data.clients.get(client_hash)
|
||||||
|
if not client or not client.ip:
|
||||||
return
|
return
|
||||||
host = await _lookup_host(ip)
|
host = await _lookup_host(client.ip)
|
||||||
country = await _geoip_country(ip)
|
country = await _geoip_country(client.ip)
|
||||||
analytics_store.enrich_visit(index, host=host, country=country)
|
city = await _geoip_city(client.ip)
|
||||||
|
analytics_store.enrich_client(client_hash, host=host, country=country, city=city)
|
||||||
|
|
||||||
|
|
||||||
|
def _schedule_client_enrichment(client_hashes: list[bytes]) -> None:
|
||||||
|
"""Start background host/geoip enrichment for the given client hashes."""
|
||||||
|
for client_hash in client_hashes:
|
||||||
|
asyncio.create_task(_enrich_client(client_hash))
|
||||||
|
|
||||||
|
|
||||||
|
async def _broadcast_analytics() -> None:
|
||||||
|
"""Send the current analytics snapshot to every connected WS client."""
|
||||||
|
if not _analytics_ws_clients:
|
||||||
|
return
|
||||||
|
payload = msgspec.json.encode(analytics_store.data).decode()
|
||||||
|
closed = set()
|
||||||
|
for ws in _analytics_ws_clients:
|
||||||
|
try:
|
||||||
|
await ws.send_text(payload)
|
||||||
|
except Exception:
|
||||||
|
closed.add(ws)
|
||||||
|
for ws in closed:
|
||||||
|
_analytics_ws_clients.discard(ws)
|
||||||
|
|
||||||
|
|
||||||
|
async def _debounced_analytics_broadcast() -> None:
|
||||||
|
"""Wait briefly, then broadcast the latest snapshot once."""
|
||||||
|
await asyncio.sleep(0.2)
|
||||||
|
await _broadcast_analytics()
|
||||||
|
|
||||||
|
|
||||||
|
def _schedule_analytics_broadcast() -> None:
|
||||||
|
"""Schedule a single debounced broadcast, ignoring duplicate triggers."""
|
||||||
|
global _analytics_broadcast_task
|
||||||
|
if _analytics_broadcast_task is not None and not _analytics_broadcast_task.done():
|
||||||
|
return
|
||||||
|
_analytics_broadcast_task = asyncio.get_running_loop().create_task(
|
||||||
|
_debounced_analytics_broadcast()
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
class AnalyticsPing(BaseModel):
|
class AnalyticsPing(BaseModel):
|
||||||
"""Navigation ping from pagerite.js (see docs/analytics.md)."""
|
"""Navigation ping from pagerite.js (see docs/analytics.md)."""
|
||||||
|
|
||||||
fr: str = ""
|
fr: str = ""
|
||||||
to: str
|
to: str | None = None
|
||||||
|
#: 1 from admin clients: scrub the session instead of recording it.
|
||||||
|
hide: int = 0
|
||||||
|
#: Active reading time on ``fr`` (ms), if any.
|
||||||
|
read: int = 0
|
||||||
|
|
||||||
|
|
||||||
@app.get("/_a", response_model=None)
|
@app.get("/_a", response_model=None)
|
||||||
async def analytics_page(request: Request) -> HTMLResponse:
|
async def analytics_page(request: Request) -> Response:
|
||||||
"""Render the analytics viewer as a normal site page at /_a.
|
"""Render the analytics viewer as a normal site page at /_a.
|
||||||
|
|
||||||
The page itself is public, but the data endpoint (/_api/analytics) stays
|
The page itself is public, but the data stream (/_api/ws/analytics) stays
|
||||||
admin-gated like the rest of /_api, so only authorized users see the
|
admin-gated like the rest of /_api, so only authorized users see the
|
||||||
statistics; others get the viewer with a "could not be loaded" message.
|
statistics; others get the viewer with a "could not be loaded" message.
|
||||||
"""
|
"""
|
||||||
return HTMLResponse(
|
return _html_response(
|
||||||
views.render_analytics(
|
request,
|
||||||
data.menu, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html
|
"analytics",
|
||||||
),
|
"",
|
||||||
headers={"cache-control": "no-cache"},
|
headers={"cache-control": "no-cache"},
|
||||||
|
etag=True,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
@@ -655,42 +838,58 @@ async def analytics_ping(ping: AnalyticsPing, request: Request) -> None:
|
|||||||
the response is never delayed by slow DNS or the first MMDB decompress.
|
the response is never delayed by slow DNS or the first MMDB decompress.
|
||||||
"""
|
"""
|
||||||
ip = _client_ip(request)
|
ip = _client_ip(request)
|
||||||
index = analytics_store.ping(
|
visit_index, flushed_clients = analytics_store.ping(
|
||||||
ping.fr,
|
ping.fr,
|
||||||
ping.to,
|
ping.to,
|
||||||
ip,
|
ip,
|
||||||
request.headers.get("user-agent", ""),
|
request.headers.get("user-agent", ""),
|
||||||
request.headers.get("accept-language", ""),
|
request.headers.get("accept-language", ""),
|
||||||
|
hide=bool(ping.hide),
|
||||||
|
read=ping.read,
|
||||||
)
|
)
|
||||||
if index is not None:
|
if visit_index is not None:
|
||||||
asyncio.create_task(_enrich_visit(index, ip))
|
visit = analytics_store.data.visits[visit_index]
|
||||||
|
asyncio.create_task(_enrich_client(visit.client))
|
||||||
|
_schedule_client_enrichment(flushed_clients)
|
||||||
|
|
||||||
|
|
||||||
def _track_entry(path: str, request: Request) -> None:
|
def _track_entry(path: str, request: Request) -> list[bytes]:
|
||||||
"""Stash the referer/UTM tags and queue a pending crawler hit for the GET.
|
"""Stash the referer/UTM tags and queue a pending crawler hit for the GET.
|
||||||
|
|
||||||
Nothing is counted on the GET itself — the client's /_a ping starts the
|
Nothing is counted on the GET itself — the client's /_a ping starts the
|
||||||
visit, so bots and admin browsing never register as visits.
|
visit, so bots never register as visits (JS-running crawlers ping too,
|
||||||
|
but the ping handler ignores known bot UAs). (Admin clients ping too,
|
||||||
|
but with hide=1, which scrubs their session instead of recording it.)
|
||||||
|
|
||||||
The devserver's health probe (``GET /?from=devserver.py`` from
|
The devserver's health probe (``GET /?from=devserver.py`` from
|
||||||
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be
|
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be
|
||||||
logged as a crawler hit. The root-path and localhost checks prevent
|
logged as a crawler hit. The root-path and localhost checks prevent
|
||||||
remote visitors from hiding traffic with the same query string.
|
remote visitors from hiding traffic with the same query string.
|
||||||
|
|
||||||
|
Returns the client hashes of any pending crawler hits flushed to persistent
|
||||||
|
storage, so callers can schedule async geoip and reverse-DNS enrichment.
|
||||||
"""
|
"""
|
||||||
|
if request.headers.get("x-pagerite-preload"):
|
||||||
|
# Idle-time page-cache warm-up by pagerite.js, not a page view: the
|
||||||
|
# ping sent when the user actually navigates does the counting.
|
||||||
|
# (Forging the header only hides a GET from the crawler stats; the
|
||||||
|
# path-based abuse classification is unaffected.)
|
||||||
|
return []
|
||||||
if (
|
if (
|
||||||
path == ""
|
path == ""
|
||||||
and str(request.url.query) == "from=devserver.py"
|
and str(request.url.query) == "from=devserver.py"
|
||||||
and _client_ip(request) == "127.0.0.1"
|
and _client_ip(request) == "127.0.0.1"
|
||||||
):
|
):
|
||||||
return
|
return []
|
||||||
own_origin = f"https://{urlparse(str(request.base_url)).netloc}"
|
own_origin = f"https://{urlparse(str(request.base_url)).netloc}"
|
||||||
analytics_store.track_entry(
|
full_path = f"{request.url.path}{_query_suffix(request)}"
|
||||||
|
return analytics_store.track_entry(
|
||||||
request.headers.get("referer", ""),
|
request.headers.get("referer", ""),
|
||||||
own_origin,
|
own_origin,
|
||||||
_client_ip(request),
|
_client_ip(request),
|
||||||
request.headers.get("user-agent", ""),
|
request.headers.get("user-agent", ""),
|
||||||
"/" if path == "" else f"/{path}",
|
full_path,
|
||||||
str(request.url.query),
|
request.headers.get("accept-language", ""),
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
@@ -727,16 +926,23 @@ def _check_reserved(path: str) -> None:
|
|||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
@app.get("/_api/analytics")
|
@app.websocket("/_api/ws/analytics")
|
||||||
async def get_analytics() -> Response:
|
async def analytics_websocket(ws: WebSocket) -> None:
|
||||||
"""The collected visit analytics as JSON (see docs/analytics.md).
|
"""Stream the analytics snapshot, then push updates as they happen.
|
||||||
|
|
||||||
Admin-only via the /_api forward-auth gate, like every management
|
Admin-only via the /_api forward-auth gate, like every management
|
||||||
endpoint. Powers the analytics viewer rendered at /_a.
|
endpoint. Powers the analytics viewer rendered at /_a.
|
||||||
"""
|
"""
|
||||||
return Response(
|
await ws.accept()
|
||||||
msgspec.json.encode(analytics_store.data), media_type="application/json"
|
await ws.send_text(msgspec.json.encode(analytics_store.data).decode())
|
||||||
)
|
_analytics_ws_clients.add(ws)
|
||||||
|
try:
|
||||||
|
while True:
|
||||||
|
await ws.receive_text()
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
finally:
|
||||||
|
_analytics_ws_clients.discard(ws)
|
||||||
|
|
||||||
|
|
||||||
@app.websocket("/_api/ws/editor")
|
@app.websocket("/_api/ws/editor")
|
||||||
@@ -890,22 +1096,108 @@ async def front_page(request: Request) -> Response:
|
|||||||
return await show_page(request, "")
|
return await show_page(request, "")
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/sitemap.xml")
|
||||||
|
async def sitemap(request: Request) -> Response:
|
||||||
|
"""Dynamically generate a sitemap of all published article pages."""
|
||||||
|
base = str(request.base_url).rstrip("/")
|
||||||
|
entries: list[tuple[str, datetime, int]] = []
|
||||||
|
|
||||||
|
def walk(
|
||||||
|
nodes: dict[str, Node], prefix: str, parent_has_content: bool = True
|
||||||
|
) -> None:
|
||||||
|
first_content_slug = next(
|
||||||
|
(
|
||||||
|
slug
|
||||||
|
for slug, node in sorted_nodes(nodes)
|
||||||
|
if node.published and node.content is not None
|
||||||
|
),
|
||||||
|
None,
|
||||||
|
)
|
||||||
|
for slug, node in sorted_nodes(nodes):
|
||||||
|
path = f"{prefix}/{slug}" if prefix else slug
|
||||||
|
depth = path.count("/") if path else 0
|
||||||
|
if (
|
||||||
|
not parent_has_content
|
||||||
|
and slug == first_content_slug
|
||||||
|
and node.published
|
||||||
|
and node.content is not None
|
||||||
|
and depth > 0
|
||||||
|
):
|
||||||
|
depth -= 1
|
||||||
|
if node.published and node.content is not None:
|
||||||
|
entries.append((path, node.modified, depth))
|
||||||
|
if node.children:
|
||||||
|
walk(node.children, path, node.content is not None)
|
||||||
|
|
||||||
|
walk(data.menu, "")
|
||||||
|
|
||||||
|
def priority(depth: int) -> float:
|
||||||
|
return max(0.1, 1.0 - depth * 0.2)
|
||||||
|
|
||||||
|
lines = [
|
||||||
|
'<?xml version="1.0" encoding="UTF-8"?>',
|
||||||
|
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
|
||||||
|
]
|
||||||
|
for path, modified, depth in entries:
|
||||||
|
loc = xml_escape(f"{base}/{path}" if path else base)
|
||||||
|
lastmod = (
|
||||||
|
modified.astimezone(UTC).replace(microsecond=0).isoformat().replace("+00:00", "Z")
|
||||||
|
)
|
||||||
|
lines.append(
|
||||||
|
f" <url>"
|
||||||
|
f"<loc>{loc}</loc>"
|
||||||
|
f"<lastmod>{lastmod}</lastmod>"
|
||||||
|
f"<priority>{priority(depth):.1f}</priority>"
|
||||||
|
f"</url>"
|
||||||
|
)
|
||||||
|
lines.append("</urlset>")
|
||||||
|
|
||||||
|
return Response(
|
||||||
|
"\n".join(lines),
|
||||||
|
media_type="application/xml",
|
||||||
|
headers={"cache-control": "no-cache"},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/robots.txt")
|
||||||
|
async def robots_txt(request: Request) -> Response:
|
||||||
|
"""Allow all crawling and point crawlers at the sitemap."""
|
||||||
|
base = str(request.base_url).rstrip("/")
|
||||||
|
body = f"User-agent: *\nAllow: /\nSitemap: {base}/sitemap.xml\n"
|
||||||
|
return Response(
|
||||||
|
body,
|
||||||
|
media_type="text/plain",
|
||||||
|
headers={"cache-control": "no-cache"},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
# Vue build asset routes are inserted at this position during load(): the
|
# Vue build asset routes are inserted at this position during load(): the
|
||||||
# build mirrors the URL space (/_assets/*, /favicon.ico at the root).
|
# build mirrors the URL space (/_assets/*, /favicon.ico at the root).
|
||||||
frontend.route(app, "/")
|
frontend.route(app, "/")
|
||||||
|
|
||||||
|
|
||||||
@app.get("/{path:path}", response_model=None)
|
@app.get("/{path:path}", response_model=None)
|
||||||
async def show_page(request: Request, path: str) -> HTMLResponse | Response:
|
async def show_page(request: Request, path: str) -> Response:
|
||||||
"""Render the content page at a slug path, or 404.
|
"""Render the content page at a slug path, or 404.
|
||||||
|
|
||||||
A node without content is a category label: its URL renders a
|
A node without content is a category label: its URL renders a
|
||||||
placeholder page (nav links point straight at its first child).
|
placeholder page (nav links point straight at its first child).
|
||||||
"""
|
"""
|
||||||
path = path.strip("/")
|
path = path.strip("/")
|
||||||
|
ua = request.headers.get("user-agent", "")
|
||||||
|
accept_language = request.headers.get("accept-language", "")
|
||||||
if path and _is_reserved(path):
|
if path and _is_reserved(path):
|
||||||
# Invalid slug shape: not a content URL, let FastAPI return its
|
# Invalid slug shape: not a content URL, let FastAPI return its
|
||||||
# built-in 404 instead of rendering an editable article page.
|
# built-in 404 instead of rendering an editable article page.
|
||||||
|
# Scanner telltales (dotpaths like /.env, *.php) classify the IP
|
||||||
|
# as abuse in analytics.
|
||||||
|
client_hash = analytics_store.track_404(
|
||||||
|
_client_ip(request),
|
||||||
|
ua,
|
||||||
|
f"/{path}{_query_suffix(request)}",
|
||||||
|
accept_language,
|
||||||
|
)
|
||||||
|
asyncio.create_task(_enrich_client(client_hash))
|
||||||
raise HTTPException(404)
|
raise HTTPException(404)
|
||||||
chain = resolve(data.menu, path)
|
chain = resolve(data.menu, path)
|
||||||
node = chain[-1] if chain else None
|
node = chain[-1] if chain else None
|
||||||
@@ -920,9 +1212,12 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
|
|||||||
if request.headers.get("if-none-match") == etag:
|
if request.headers.get("if-none-match") == etag:
|
||||||
return Response(status_code=304)
|
return Response(status_code=304)
|
||||||
if _is_trackable_path(path):
|
if _is_trackable_path(path):
|
||||||
_track_entry(path, request)
|
flushed = _track_entry(path, request)
|
||||||
return HTMLResponse(
|
_schedule_client_enrichment(flushed)
|
||||||
views.render_page(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html, str(request.base_url).rstrip("/")),
|
return _html_response(
|
||||||
|
request,
|
||||||
|
"page",
|
||||||
|
path,
|
||||||
headers={
|
headers={
|
||||||
"etag": etag,
|
"etag": etag,
|
||||||
"last-modified": _http_date(node.modified),
|
"last-modified": _http_date(node.modified),
|
||||||
@@ -933,9 +1228,12 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
|
|||||||
# Category label without a landing page: placeholder with the pen
|
# Category label without a landing page: placeholder with the pen
|
||||||
# to create it (404 — no page here, but the node is real).
|
# to create it (404 — no page here, but the node is real).
|
||||||
if _is_trackable_path(path):
|
if _is_trackable_path(path):
|
||||||
_track_entry(path, request)
|
flushed = _track_entry(path, request)
|
||||||
return HTMLResponse(
|
_schedule_client_enrichment(flushed)
|
||||||
views.render_category(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html),
|
return _html_response(
|
||||||
|
request,
|
||||||
|
"category",
|
||||||
|
path,
|
||||||
404,
|
404,
|
||||||
headers={
|
headers={
|
||||||
"last-modified": _http_date(node.modified),
|
"last-modified": _http_date(node.modified),
|
||||||
@@ -949,5 +1247,13 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
|
|||||||
if item.published:
|
if item.published:
|
||||||
return RedirectResponse(f"/{slug}")
|
return RedirectResponse(f"/{slug}")
|
||||||
if _is_trackable_path(path):
|
if _is_trackable_path(path):
|
||||||
_track_entry(path, request)
|
client_hash = analytics_store.track_404(
|
||||||
return HTMLResponse(views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html), 404)
|
_client_ip(request),
|
||||||
|
ua,
|
||||||
|
f"/{path}{_query_suffix(request)}",
|
||||||
|
accept_language,
|
||||||
|
)
|
||||||
|
asyncio.create_task(_enrich_client(client_hash))
|
||||||
|
flushed = _track_entry(path, request)
|
||||||
|
_schedule_client_enrichment(flushed)
|
||||||
|
return _html_response(request, "not-found", path, 404)
|
||||||
|
|||||||
@@ -100,6 +100,12 @@ class Data(msgspec.Struct):
|
|||||||
#: Favicon: name of a file in `files` (content-addressed), linked as
|
#: Favicon: name of a file in `files` (content-addressed), linked as
|
||||||
#: <link rel="icon"> on every page. Empty = the build's /favicon.ico.
|
#: <link rel="icon"> on every page. Empty = the build's /favicon.ico.
|
||||||
favicon: str = ""
|
favicon: str = ""
|
||||||
|
#: Public origin (scheme + host) of the site, learned from admin
|
||||||
|
#: browsers (POST /_api/site-url — location.origin is correct even
|
||||||
|
#: behind reverse proxies, unlike request Host headers). Used for
|
||||||
|
#: absolute social/canonical URLs; empty = fall back to the request's
|
||||||
|
#: own base URL.
|
||||||
|
site_url: str = ""
|
||||||
#: Legacy flat page store (pre-tree databases); migrated into `menu`
|
#: Legacy flat page store (pre-tree databases); migrated into `menu`
|
||||||
#: on startup, then cleared. Never written otherwise.
|
#: on startup, then cleared. Never written otherwise.
|
||||||
pages: dict[str, Page] = {}
|
pages: dict[str, Page] = {}
|
||||||
|
|||||||
+5
-3
@@ -29,6 +29,7 @@ Where to go next:
|
|||||||
- The [docs](/docs/editing) section explains how to edit this site and shows every supported Markdown feature, source and result side by side.
|
- The [docs](/docs/editing) section explains how to edit this site and shows every supported Markdown feature, source and result side by side.
|
||||||
- The [showcase](/showcase/gallery) section shows what finished pages can look like: image positioning, banners, a long read.
|
- The [showcase](/showcase/gallery) section shows what finished pages can look like: image positioning, banners, a long read.
|
||||||
- Click the 🖊️ pen on any page to open the editor, and the ⚙️ pen for site settings and the structure tree.
|
- Click the 🖊️ pen on any page to open the editor, and the ⚙️ pen for site settings and the structure tree.
|
||||||
|
- Elsewhere on the web: [{width=240}](https://xkcd.com/927/) — a cautionary tale about adding one more standard.
|
||||||
|
|
||||||
{width=420}
|
{width=420}
|
||||||
|
|
||||||
@@ -68,15 +69,15 @@ Every feature below is shown twice: first the Markdown source, then how it rende
|
|||||||
### A subsection
|
### A subsection
|
||||||
|
|
||||||
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a
|
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a
|
||||||
[link to the front page](/). Plain URLs become links automatically:
|
[link to the front page](/). An image that links to its page:
|
||||||
https://example.com — and a hard line break
|
[{width=240}](https://xkcd.com/1179/) — and a hard line break
|
||||||
is just a newline.
|
is just a newline.
|
||||||
```
|
```
|
||||||
|
|
||||||
## A section heading
|
## A section heading
|
||||||
### A subsection
|
### A subsection
|
||||||
|
|
||||||
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a [link to the front page](/). Plain URLs become links automatically: https://example.com — and a hard line break
|
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a [link to the front page](/). An image that links to its page: [{width=240}](https://xkcd.com/1179/) — and a hard line break
|
||||||
is just a newline.
|
is just a newline.
|
||||||
|
|
||||||
## Lists and quotes
|
## Lists and quotes
|
||||||
@@ -354,6 +355,7 @@ This site runs on **Pagerite**: FastAPI + html5tagger + kanta, with content writ
|
|||||||
- [How to edit this site](/docs/editing)
|
- [How to edit this site](/docs/editing)
|
||||||
- [Markdown features](/docs/markdown/basics)
|
- [Markdown features](/docs/markdown/basics)
|
||||||
- [The showcase](/showcase/gallery)
|
- [The showcase](/showcase/gallery)
|
||||||
|
- [{width=240}](https://xkcd.com/2347/) — a small comic about small dependencies
|
||||||
|
|
||||||
*Replace this page with whatever your site is about.*
|
*Replace this page with whatever your site is about.*
|
||||||
"""
|
"""
|
||||||
|
|||||||
+114
-30
@@ -110,6 +110,35 @@ def _editor_css_url(vite_url: str | None) -> str | None:
|
|||||||
return None
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _inline_asset(url: str) -> str:
|
||||||
|
"""Read a served asset's content for inlining into the page (prod only).
|
||||||
|
|
||||||
|
Handles build assets (``/_assets/...`` from the Vite build) and theme
|
||||||
|
files (``/_themes/{name}/...`` from pagerite/themes/).
|
||||||
|
"""
|
||||||
|
if url.startswith("/_themes/"):
|
||||||
|
name, _, file = url.removeprefix("/_themes/").partition("/")
|
||||||
|
if _valid_name(name) and _valid_name(file):
|
||||||
|
return (THEMES / name / file).read_text()
|
||||||
|
raise ValueError(f"not a theme asset: {url}")
|
||||||
|
return (BUILD / url.lstrip("/")).read_text()
|
||||||
|
|
||||||
|
|
||||||
|
def _inline_script(url: str) -> str:
|
||||||
|
"""Read a built JS bundle for inlining (prod only).
|
||||||
|
|
||||||
|
Inline modules resolve relative imports against the document URL, not
|
||||||
|
the bundle's directory, so rewrite the build's relative chunk
|
||||||
|
specifiers ("./chunk.js") to absolute /_assets/ paths.
|
||||||
|
"""
|
||||||
|
js = _inline_asset(url)
|
||||||
|
for chunk in _manifest().values():
|
||||||
|
file = chunk.get("file", "")
|
||||||
|
if file.endswith(".js"):
|
||||||
|
js = js.replace(f'"./{file.rsplit("/", 1)[-1]}"', f'"/{file}"')
|
||||||
|
return js
|
||||||
|
|
||||||
|
|
||||||
def _layout(
|
def _layout(
|
||||||
modules: list[str] = (),
|
modules: list[str] = (),
|
||||||
stylesheets: list[str] = (),
|
stylesheets: list[str] = (),
|
||||||
@@ -118,16 +147,21 @@ def _layout(
|
|||||||
banner_design: str = "",
|
banner_design: str = "",
|
||||||
favicon: str = "",
|
favicon: str = "",
|
||||||
social: dict[str, str] | None = None,
|
social: dict[str, str] | None = None,
|
||||||
extra_meta: dict[str, str] | None = None,
|
|
||||||
) -> Template:
|
) -> Template:
|
||||||
"""Page layout template with standard asset URLs and ES-module scripts.
|
"""Page layout template with standard assets and ES-module scripts.
|
||||||
|
|
||||||
Stylesheets use ``blocking="render"`` so the browser waits for them before
|
In dev (PAGERITE_VITE_URL set) assets are linked from the Vite dev
|
||||||
showing the page, avoiding a flash of unstyled content. Order matters and
|
server and stylesheets use ``blocking="render"`` so the browser waits
|
||||||
is fixed: base (Vite build, absent in dev where Vite injects it from JS),
|
for them before showing the page, avoiding a flash of unstyled content.
|
||||||
theme and banner design (backend-served from pagerite/themes/), entry-
|
In production all page assets are inlined into the document: stylesheets
|
||||||
specific stylesheets (e.g. overlayscrollbars.css), then the user's custom
|
become ``<style>`` elements and module scripts inline ``<script>``s, so
|
||||||
CSS last so it always wins.
|
a page loads with no asset round trips. The on-demand bundles (editor,
|
||||||
|
analytics) stay external in both modes.
|
||||||
|
|
||||||
|
Order matters and is fixed: base (Vite build, absent in dev where Vite
|
||||||
|
injects it from JS), theme and banner design (from pagerite/themes/),
|
||||||
|
entry-specific stylesheets (e.g. overlayscrollbars.css), then the user's
|
||||||
|
custom CSS last so it always wins.
|
||||||
|
|
||||||
In dev, pagerite.js re-appends the backend-rendered theme/design links
|
In dev, pagerite.js re-appends the backend-rendered theme/design links
|
||||||
(and the custom CSS) after the Vite-injected base styles, keeping this
|
(and the custom CSS) after the Vite-injected base styles, keeping this
|
||||||
@@ -135,9 +169,6 @@ def _layout(
|
|||||||
|
|
||||||
``social`` maps meta keys to contents: ``og:*``/``article:*`` go out as
|
``social`` maps meta keys to contents: ``og:*``/``article:*`` go out as
|
||||||
property attributes, everything else (description, twitter:*) as name.
|
property attributes, everything else (description, twitter:*) as name.
|
||||||
|
|
||||||
``extra_meta`` is emitted as plain ``<meta name="..." content="...">``
|
|
||||||
tags after the editor meta tags; used for page-specific import hints.
|
|
||||||
"""
|
"""
|
||||||
doc = Document(E.Title, lang="en")
|
doc = Document(E.Title, lang="en")
|
||||||
# Responsive layout (see the 48rem breakpoint in pagerite.css) needs
|
# Responsive layout (see the 48rem breakpoint in pagerite.css) needs
|
||||||
@@ -155,33 +186,63 @@ def _layout(
|
|||||||
# one, browsers fall back to the build's /favicon.ico by convention.
|
# one, browsers fall back to the build's /favicon.ico by convention.
|
||||||
if favicon:
|
if favicon:
|
||||||
doc.link(rel="icon", href=f"/_f/{favicon}", id="pagerite-favicon")
|
doc.link(rel="icon", href=f"/_f/{favicon}", id="pagerite-favicon")
|
||||||
# Editor asset URLs for pagerite.js, which injects the 🖊️ edit pens
|
# Asset URLs for the on-demand bundles (editor, analytics) for
|
||||||
# itself once it has validated the session (pages render identically
|
# pagerite.js, which injects the 🖊️ edit pens itself once it has
|
||||||
# for everyone; editing is gated by the auth proxy in front of /_api).
|
# validated the session (pages render identically for everyone; editing
|
||||||
script, editor_css = _editor_assets()
|
# is gated by the auth proxy in front of /_api). Dev passes the Vite
|
||||||
doc.meta(name="pagerite:editor-src", content=script[-1])
|
# dev-server URLs as meta tags (Vite serves the modules and injects
|
||||||
if editor_css:
|
# their CSS for hot reloads); production inlines all page assets and
|
||||||
doc.meta(name="pagerite:editor-css", content=editor_css)
|
# carries the on-demand URLs in one JSON script instead.
|
||||||
for key, value in (extra_meta or {}).items():
|
|
||||||
doc.meta(name=key, content=value)
|
|
||||||
# Stylesheet links carry stable ids so the site editor's hot swap can
|
|
||||||
# keep each sheet at its rendered position (see swapRegions).
|
|
||||||
vite_url = os.environ.get("PAGERITE_VITE_URL")
|
vite_url = os.environ.get("PAGERITE_VITE_URL")
|
||||||
|
editor_scripts, editor_css = _editor_assets()
|
||||||
|
config = {
|
||||||
|
"pagerite:editor-src": editor_scripts[-1],
|
||||||
|
"pagerite:analytics-src": _analytics_assets()[0][0],
|
||||||
|
}
|
||||||
|
if editor_css:
|
||||||
|
config["pagerite:editor-css"] = editor_css
|
||||||
|
if vite_url:
|
||||||
|
for key, value in config.items():
|
||||||
|
doc.meta(name=key, content=value)
|
||||||
|
else:
|
||||||
|
# Inert JSON script; URLs never contain "</", but stay safe.
|
||||||
|
doc.script(
|
||||||
|
HTML(json.dumps(config).replace("</", "<\\/")),
|
||||||
|
type="application/json",
|
||||||
|
id="pagerite-assets",
|
||||||
|
)
|
||||||
|
# Stylesheets carry stable ids so the fetch-navigation and the site
|
||||||
|
# editor's hot swap can sync <head> positionally (see swapdoc.js).
|
||||||
|
# Production inlines the CSS as <style> elements: one less round trip
|
||||||
|
# per sheet, and fetch-navigation can carry them across swaps whole.
|
||||||
sheets = [
|
sheets = [
|
||||||
("pagerite-base", _base_css_url(vite_url)),
|
("pagerite-base", _base_css_url(vite_url)),
|
||||||
("pagerite-theme", _theme_css_url(theme)),
|
("pagerite-theme", _theme_css_url(theme)),
|
||||||
("pagerite-banner", _banner_css_url(banner_design)),
|
("pagerite-banner", _banner_css_url(banner_design)),
|
||||||
]
|
]
|
||||||
for id_, url in sheets:
|
for id_, url in sheets:
|
||||||
if url:
|
if not url:
|
||||||
|
continue
|
||||||
|
if vite_url:
|
||||||
doc.link(rel="stylesheet", href=url, blocking="render", id=id_)
|
doc.link(rel="stylesheet", href=url, blocking="render", id=id_)
|
||||||
|
else:
|
||||||
|
doc.style(HTML(_inline_asset(url)), id=id_)
|
||||||
for url in stylesheets:
|
for url in stylesheets:
|
||||||
doc.link(rel="stylesheet", href=url, blocking="render")
|
if vite_url:
|
||||||
|
doc.link(rel="stylesheet", href=url, blocking="render")
|
||||||
|
else:
|
||||||
|
# Id from the file stem minus the content hash, so the head
|
||||||
|
# sync can match sheets across pages (e.g. the analytics sheet
|
||||||
|
# exists on /_a only and is added/removed on swaps).
|
||||||
|
stem = url.rsplit("/", 1)[-1].removesuffix(".css")
|
||||||
|
name = re.sub(r"-[A-Za-z0-9_-]{8}$", "", stem)
|
||||||
|
doc.style(HTML(_inline_asset(url)), id=f"pagerite-css-{name}")
|
||||||
for src in modules:
|
for src in modules:
|
||||||
doc.script(src=src, type="module")
|
if vite_url:
|
||||||
|
doc.script(src=src, type="module")
|
||||||
if custom_css.strip():
|
if custom_css.strip():
|
||||||
doc.style(custom_css, id="pagerite-user")
|
doc.style(custom_css, id="pagerite-user")
|
||||||
return Template(
|
body = (
|
||||||
doc
|
doc
|
||||||
.header(
|
.header(
|
||||||
E.div(E.Banner, id="page-banner"),
|
E.div(E.Banner, id="page-banner"),
|
||||||
@@ -194,8 +255,23 @@ def _layout(
|
|||||||
E.main(E.Main, id="main"),
|
E.main(E.Main, id="main"),
|
||||||
id="content",
|
id="content",
|
||||||
)
|
)
|
||||||
.footer(None), # kept empty for now; zero-height (see pagerite.css)
|
.footer(None) # kept empty for now; zero-height (see pagerite.css)
|
||||||
)
|
)
|
||||||
|
if not vite_url:
|
||||||
|
# Inline the bundles at the end of the body: module scripts are
|
||||||
|
# deferred anyway, and the page can render before they execute.
|
||||||
|
# Escape "</script" so it cannot terminate the element early (only
|
||||||
|
# ever occurs inside string literals, where the backslash escape is
|
||||||
|
# a no-op).
|
||||||
|
for src in modules:
|
||||||
|
js = re.sub(r"</script", r"<\\/script", _inline_script(src), flags=re.I)
|
||||||
|
# Stable id from the file stem minus the content hash; the
|
||||||
|
# analytics page's script (pagerite-js-analytics) is found and
|
||||||
|
# re-created by pagerite.js on fetch-navigations to /_a.
|
||||||
|
stem = src.rsplit("/", 1)[-1].removesuffix(".js")
|
||||||
|
name = re.sub(r"-[A-Za-z0-9_-]{8}$", "", stem)
|
||||||
|
body.script(HTML(js), type="module", id=f"pagerite-js-{name}")
|
||||||
|
return Template(body)
|
||||||
|
|
||||||
|
|
||||||
def _brand_link(brand: str, brand_html: str = "") -> HTML:
|
def _brand_link(brand: str, brand_html: str = "") -> HTML:
|
||||||
@@ -682,14 +758,23 @@ def render_analytics(
|
|||||||
favicon: str = "",
|
favicon: str = "",
|
||||||
brand_html: str = "",
|
brand_html: str = "",
|
||||||
) -> str:
|
) -> str:
|
||||||
"""Render the analytics viewer as a normal page at /_a."""
|
"""Render the analytics viewer as a normal page at /_a.
|
||||||
|
|
||||||
|
The analytics entry is inlined into this page only (prod) or loaded
|
||||||
|
from the Vite dev server (dev); its stylesheet rides along in <head>
|
||||||
|
so fetch-navigations can sync it into the live document. The initial
|
||||||
|
range is not rendered in: the client takes it from the URL hash or
|
||||||
|
derives it from the analytics data itself.
|
||||||
|
"""
|
||||||
page_scripts, page_stylesheets = _page_assets()
|
page_scripts, page_stylesheets = _page_assets()
|
||||||
analytics_scripts, analytics_stylesheets = _analytics_assets()
|
analytics_scripts, analytics_stylesheets = _analytics_assets()
|
||||||
scripts = page_scripts + analytics_scripts
|
scripts = page_scripts + analytics_scripts
|
||||||
stylesheets = page_stylesheets + analytics_stylesheets
|
stylesheets = page_stylesheets + analytics_stylesheets
|
||||||
doc = E.article
|
doc = E.article
|
||||||
with doc:
|
with doc:
|
||||||
doc.div(id="analytics-app")
|
# .wide: the dashboard breaks out of the article column to the full
|
||||||
|
# viewport width, like wide figures (see the .wide rules).
|
||||||
|
doc.div(id="analytics-app", class_="wide")
|
||||||
return str(
|
return str(
|
||||||
_layout(
|
_layout(
|
||||||
scripts,
|
scripts,
|
||||||
@@ -698,7 +783,6 @@ def render_analytics(
|
|||||||
theme,
|
theme,
|
||||||
banner_design(menu, "_a", theme),
|
banner_design(menu, "_a", theme),
|
||||||
favicon,
|
favicon,
|
||||||
extra_meta={"pagerite:analytics-src": analytics_scripts[0]},
|
|
||||||
)(
|
)(
|
||||||
Title=f"Analytics – {brand}" if brand else "Analytics",
|
Title=f"Analytics – {brand}" if brand else "Analytics",
|
||||||
Brand=_brand_link(brand, brand_html),
|
Brand=_brand_link(brand, brand_html),
|
||||||
|
|||||||
+3
-3
@@ -20,12 +20,14 @@ dependencies = [
|
|||||||
"fastapi-vue>=1.3.1",
|
"fastapi-vue>=1.3.1",
|
||||||
"fastapi[standard]>=0.141.1",
|
"fastapi[standard]>=0.141.1",
|
||||||
"html5tagger>=2.0.0",
|
"html5tagger>=2.0.0",
|
||||||
|
"httpx>=0.28.1",
|
||||||
"kanta>=0.8.1",
|
"kanta>=0.8.1",
|
||||||
"markdown-it-py>=4.2.0",
|
"markdown-it-py>=4.2.0",
|
||||||
"maxminddb>=3.1.1",
|
"maxminddb>=3.1.1",
|
||||||
"mdit-py-plugins>=0.6.1",
|
"mdit-py-plugins>=0.6.1",
|
||||||
"pygments>=2.20.0",
|
"pygments>=2.20.0",
|
||||||
"ua-parser>=1.0.2",
|
"ua-parser>=1.0.2",
|
||||||
|
"zstandard>=0.25.0",
|
||||||
]
|
]
|
||||||
|
|
||||||
[project.scripts]
|
[project.scripts]
|
||||||
@@ -35,9 +37,7 @@ pagerite = "pagerite.__main__:main"
|
|||||||
Repository = "https://git.zi.fi/LeoVasanko/pagerite"
|
Repository = "https://git.zi.fi/LeoVasanko/pagerite"
|
||||||
|
|
||||||
[dependency-groups]
|
[dependency-groups]
|
||||||
dev = [
|
dev = []
|
||||||
"httpx>=0.28.1",
|
|
||||||
]
|
|
||||||
|
|
||||||
[tool.hatch.version]
|
[tool.hatch.version]
|
||||||
source = "vcs"
|
source = "vcs"
|
||||||
|
|||||||
@@ -19,8 +19,8 @@ from devutil import (
|
|||||||
setup_vite,
|
setup_vite,
|
||||||
)
|
)
|
||||||
|
|
||||||
DEFAULT_VITE_PORT = 3100
|
DEFAULT_VITE_PORT = 8200
|
||||||
DEFAULT_DEV_PORT = 3200
|
DEFAULT_DEV_PORT = 8210
|
||||||
HEALTH = "/?from=devserver.py"
|
HEALTH = "/?from=devserver.py"
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
+664
-143
@@ -6,20 +6,18 @@
|
|||||||
# "playwright>=1.45.0",
|
# "playwright>=1.45.0",
|
||||||
# ]
|
# ]
|
||||||
# ///
|
# ///
|
||||||
"""Generate fake browser visits and crawler hits for a Pagerite site.
|
"""Generate fake browser visits, crawler hits, and abuse scans for a Pagerite site.
|
||||||
|
|
||||||
The script drives a real Chromium browser with Playwright, clicking visible
|
Browser sessions (ordinary users) come from realistic residential IPv4 and IPv6
|
||||||
internal links so the site's own analytics JavaScript records normal visits
|
addresses and stay mostly stable; an IPv6 host part may rotate once mid-session,
|
||||||
(POST /_a). Browser sessions and crawler GETs send a small rotating pool of
|
and an IPv4 session may switch to another residential address. Crawler hits come
|
||||||
real public IPs in X-Forwarded-For, so the backend can reverse-DNS and GeoIP
|
from datacenter IPs, with each crawler profile paired to a matching provider IP
|
||||||
them instead of seeing every hit as 127.0.0.1.
|
when possible. Abuse scanners fire bursts of vulnerability probes from pinned
|
||||||
|
datacenter IPs.
|
||||||
Sessions start with a Poisson inter-arrival delay (``--arrival-rate``) to
|
|
||||||
spread traffic out a little, while still keeping the overall run fast.
|
|
||||||
|
|
||||||
Run against a local dev server, e.g.:
|
Run against a local dev server, e.g.:
|
||||||
|
|
||||||
uv run scripts/fake_traffic.py http://localhost:3200 -b 8 -c 20
|
uv run scripts/fake_traffic.py http://localhost:3200
|
||||||
|
|
||||||
Repeat whenever you want more traffic; each run appends new events to the
|
Repeat whenever you want more traffic; each run appends new events to the
|
||||||
site's analytics file.
|
site's analytics file.
|
||||||
@@ -36,7 +34,7 @@ from collections.abc import Sequence
|
|||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from datetime import UTC, datetime
|
from datetime import UTC, datetime
|
||||||
from typing import Any
|
from typing import Any
|
||||||
from urllib.parse import urljoin, urlparse
|
from urllib.parse import urlencode, urljoin, urlparse
|
||||||
|
|
||||||
import httpx
|
import httpx
|
||||||
|
|
||||||
@@ -56,6 +54,7 @@ class BrowserProfile:
|
|||||||
class CrawlerProfile:
|
class CrawlerProfile:
|
||||||
name: str
|
name: str
|
||||||
user_agent: str
|
user_agent: str
|
||||||
|
ip: str
|
||||||
|
|
||||||
|
|
||||||
BROWSER_PROFILES: list[BrowserProfile] = [
|
BROWSER_PROFILES: list[BrowserProfile] = [
|
||||||
@@ -92,37 +91,419 @@ CRAWLER_PROFILES: list[CrawlerProfile] = [
|
|||||||
CrawlerProfile(
|
CrawlerProfile(
|
||||||
"googlebot",
|
"googlebot",
|
||||||
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
|
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
|
||||||
|
"66.249.64.66", # US, Google
|
||||||
),
|
),
|
||||||
CrawlerProfile(
|
CrawlerProfile(
|
||||||
"bingbot",
|
"bingbot",
|
||||||
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
|
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
|
||||||
|
"40.77.167.0", # US, Microsoft
|
||||||
),
|
),
|
||||||
CrawlerProfile(
|
CrawlerProfile(
|
||||||
"duckduckbot", "DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)"
|
"duckduckbot",
|
||||||
|
"DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)",
|
||||||
|
"95.217.0.1", # Germany, Hetzner VPS
|
||||||
|
),
|
||||||
|
CrawlerProfile(
|
||||||
|
"curl",
|
||||||
|
"curl/8.5.0",
|
||||||
|
"139.162.0.1", # Singapore, Linode VPS
|
||||||
),
|
),
|
||||||
CrawlerProfile("curl", "curl/8.5.0"),
|
|
||||||
]
|
]
|
||||||
|
|
||||||
# Small pool of real public resolver IPs. They have real reverse-DNS and GeoIP
|
# Residential IPv4 addresses and IPv6 /64 prefixes used for ordinary browser
|
||||||
# entries, and cycling through a handful avoids hammering DNS during traffic
|
# sessions. IPv6 entries keep the network part stable and randomise only the
|
||||||
# generation.
|
# host part; the host may rotate once mid-session.
|
||||||
SOURCE_IPS: list[str] = [
|
RESIDENTIAL_SOURCE_IPS: list[str] = [
|
||||||
"8.8.8.8",
|
# Residential IPv4
|
||||||
"1.1.1.1",
|
"91.154.140.209", # Finland, Elisa
|
||||||
"9.9.9.9",
|
"84.143.145.207", # Germany, Deutsche Telekom
|
||||||
"208.67.222.222",
|
"220.165.255.254", # China, Chinanet / China Telecom
|
||||||
"185.228.168.9",
|
"84.235.83.162", # Saudi Arabia, SaudiNet / STC
|
||||||
"94.140.14.14",
|
# Residential IPv6 /64 prefixes
|
||||||
|
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
|
||||||
|
"240e:45d:1e60:5b0::/64", # China, China Telecom
|
||||||
|
"2409:8904:6720:4123::/64", # China, China Unicom
|
||||||
]
|
]
|
||||||
|
|
||||||
|
# Concrete datacenter IPs used for abuse scanner bursts. They stay pinned for
|
||||||
|
# the whole scan burst.
|
||||||
|
# Index 0 randomises its UA per request, index 1 uses a fixed browser UA,
|
||||||
|
# and index 2 uses a fixed crawler UA.
|
||||||
|
ABUSE_SOURCE_IPS: list[str] = [
|
||||||
|
"45.63.0.12", # US, Vultr VPS
|
||||||
|
"138.197.0.89", # US, DigitalOcean / Cloudways
|
||||||
|
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS
|
||||||
|
]
|
||||||
|
|
||||||
|
# Paths commonly probed by attackers looking for exposed config, admin panels,
|
||||||
|
# version control, credentials, backups, or debug endpoints.
|
||||||
|
SUSPICIOUS_PATHS: list[str] = [
|
||||||
|
"/.env",
|
||||||
|
"/env",
|
||||||
|
"/.env.local",
|
||||||
|
"/env.development",
|
||||||
|
"/config",
|
||||||
|
"/config.json",
|
||||||
|
"/config.yaml",
|
||||||
|
"/config.yml",
|
||||||
|
"/configuration.json",
|
||||||
|
"/configuration.yaml",
|
||||||
|
"/configuration.yml",
|
||||||
|
"/settings.json",
|
||||||
|
"/settings.yaml",
|
||||||
|
"/settings.yml",
|
||||||
|
"/app.config",
|
||||||
|
"/appsettings.json",
|
||||||
|
"/appsettings.Development.json",
|
||||||
|
"/credentials",
|
||||||
|
"/credentials.json",
|
||||||
|
"/secrets",
|
||||||
|
"/secrets.json",
|
||||||
|
"/.aws/credentials",
|
||||||
|
"/.ssh/id_rsa",
|
||||||
|
"/id_rsa",
|
||||||
|
"/id_rsa.pub",
|
||||||
|
"/known_hosts",
|
||||||
|
"/sftp-config.json",
|
||||||
|
"/admin",
|
||||||
|
"/administrator",
|
||||||
|
"/adminer.php",
|
||||||
|
"/login",
|
||||||
|
"/signin",
|
||||||
|
"/auth/login",
|
||||||
|
"/api/login",
|
||||||
|
"/api/.env",
|
||||||
|
"/api/config",
|
||||||
|
"/api/v1/config",
|
||||||
|
"/api/v2/config",
|
||||||
|
"/webhook",
|
||||||
|
"/webhooks",
|
||||||
|
"/callback",
|
||||||
|
"/proxy",
|
||||||
|
"/image",
|
||||||
|
"/images",
|
||||||
|
"/preview",
|
||||||
|
"/download",
|
||||||
|
"/downloads",
|
||||||
|
"/log",
|
||||||
|
"/logs",
|
||||||
|
"/debug",
|
||||||
|
"/trace",
|
||||||
|
"/phpinfo.php",
|
||||||
|
"/info.php",
|
||||||
|
"/phpmyadmin",
|
||||||
|
"/pma",
|
||||||
|
"/myadmin",
|
||||||
|
"/phpMyAdmin",
|
||||||
|
"/wp-admin",
|
||||||
|
"/wp-login.php",
|
||||||
|
"/wp-config.php",
|
||||||
|
"/xmlrpc.php",
|
||||||
|
"/wp-json/wp/v2/users",
|
||||||
|
"/.git/config",
|
||||||
|
"/.git/HEAD",
|
||||||
|
"/git/config",
|
||||||
|
"/swagger-ui.html",
|
||||||
|
"/v2/api-docs",
|
||||||
|
"/actuator/env",
|
||||||
|
"/actuator/health",
|
||||||
|
"/actuator/configprops",
|
||||||
|
"/server-status",
|
||||||
|
"/.htaccess",
|
||||||
|
"/web.config",
|
||||||
|
"/package.json",
|
||||||
|
"/composer.json",
|
||||||
|
"/vendor/autoload.php",
|
||||||
|
"/docker-compose.yml",
|
||||||
|
"/Dockerfile",
|
||||||
|
"/manage",
|
||||||
|
"/console",
|
||||||
|
"/manager",
|
||||||
|
"/manager/html",
|
||||||
|
"/metrics",
|
||||||
|
"/prometheus",
|
||||||
|
"/healthz",
|
||||||
|
"/_api",
|
||||||
|
"/api",
|
||||||
|
"/api/v1/",
|
||||||
|
"/api/v2/",
|
||||||
|
"/graphql",
|
||||||
|
"/query",
|
||||||
|
"/feed",
|
||||||
|
"/rss",
|
||||||
|
"/_debug",
|
||||||
|
"/test",
|
||||||
|
"/testing",
|
||||||
|
"/tmp",
|
||||||
|
"/temp",
|
||||||
|
"/backup",
|
||||||
|
"/backups",
|
||||||
|
"/dump",
|
||||||
|
"/dumps",
|
||||||
|
"/sql",
|
||||||
|
"/db",
|
||||||
|
"/database",
|
||||||
|
"/dump.sql",
|
||||||
|
"/backup.sql",
|
||||||
|
"/db.sql",
|
||||||
|
"/backup.zip",
|
||||||
|
"/backup.tar.gz",
|
||||||
|
"/site.zip",
|
||||||
|
"/site.tar.gz",
|
||||||
|
"/source.zip",
|
||||||
|
"/src.zip",
|
||||||
|
"/upload",
|
||||||
|
"/uploads",
|
||||||
|
"/import",
|
||||||
|
"/export",
|
||||||
|
"/token",
|
||||||
|
"/tokens",
|
||||||
|
"/oauth",
|
||||||
|
"/oauth2",
|
||||||
|
"/openid",
|
||||||
|
"/jwks",
|
||||||
|
"/keys",
|
||||||
|
"/key",
|
||||||
|
"/private",
|
||||||
|
"/public",
|
||||||
|
]
|
||||||
|
|
||||||
|
# Realistic external referers. Most sessions arrive with a generic referer;
|
||||||
|
# a subset carries matching UTM tags on the landing URL.
|
||||||
|
PLAIN_REFERRERS: list[str] = [
|
||||||
|
"https://example.com/",
|
||||||
|
"https://somedomain.com/",
|
||||||
|
"https://another-site.org/",
|
||||||
|
"https://friend-site.net/",
|
||||||
|
]
|
||||||
|
|
||||||
|
# (referer origin, utm parameter dict) pairs used for tagged traffic.
|
||||||
|
TAGGED_REFERRERS: list[tuple[str, dict[str, str]]] = [
|
||||||
|
("https://chatgpt.com/", {"utm_source": "chatgpt.com"}),
|
||||||
|
("https://www.google.com/", {"utm_source": "google", "utm_medium": "organic"}),
|
||||||
|
("https://twitter.com/", {"utm_source": "twitter", "utm_medium": "social"}),
|
||||||
|
("https://www.linkedin.com/", {"utm_source": "linkedin", "utm_medium": "social"}),
|
||||||
|
("https://github.com/", {"utm_source": "github", "utm_medium": "referral"}),
|
||||||
|
("https://news.ycombinator.com/", {"utm_source": "hackernews", "utm_medium": "referral"}),
|
||||||
|
("https://www.reddit.com/", {"utm_source": "reddit", "utm_medium": "social"}),
|
||||||
|
("https://medium.com/", {"utm_source": "medium", "utm_medium": "referral"}),
|
||||||
|
("https://www.producthunt.com/", {"utm_source": "producthunt", "utm_medium": "referral"}),
|
||||||
|
]
|
||||||
|
|
||||||
|
# Fraction of referered sessions that also carry UTM tags.
|
||||||
|
UTM_RATE = 0.25
|
||||||
|
|
||||||
|
# Innocent-looking paths that do not exist on a Pagerite site. Hitting many of
|
||||||
|
# these from a single IP is itself a telltale of a spray-and-pray scanner.
|
||||||
|
NORMAL_404_PATHS: list[str] = [
|
||||||
|
"/about",
|
||||||
|
"/about-us",
|
||||||
|
"/services",
|
||||||
|
"/products",
|
||||||
|
"/contact",
|
||||||
|
"/contact-us",
|
||||||
|
"/team",
|
||||||
|
"/careers",
|
||||||
|
"/jobs",
|
||||||
|
"/pricing",
|
||||||
|
"/features",
|
||||||
|
"/demo",
|
||||||
|
"/trial",
|
||||||
|
"/docs",
|
||||||
|
"/documentation",
|
||||||
|
"/api-docs",
|
||||||
|
"/support",
|
||||||
|
"/help",
|
||||||
|
"/faq",
|
||||||
|
"/knowledge-base",
|
||||||
|
"/terms",
|
||||||
|
"/terms-of-service",
|
||||||
|
"/privacy",
|
||||||
|
"/privacy-policy",
|
||||||
|
"/legal",
|
||||||
|
"/blog",
|
||||||
|
"/news",
|
||||||
|
"/articles",
|
||||||
|
"/press",
|
||||||
|
"/events",
|
||||||
|
"/webinars",
|
||||||
|
"/podcast",
|
||||||
|
"/videos",
|
||||||
|
"/resources",
|
||||||
|
"/whitepapers",
|
||||||
|
"/case-studies",
|
||||||
|
"/customers",
|
||||||
|
"/clients",
|
||||||
|
"/testimonials",
|
||||||
|
"/reviews",
|
||||||
|
"/partners",
|
||||||
|
"/integrations",
|
||||||
|
"/api-reference",
|
||||||
|
"/developers",
|
||||||
|
"/status",
|
||||||
|
"/security",
|
||||||
|
"/trust",
|
||||||
|
"/compliance",
|
||||||
|
"/gdpr",
|
||||||
|
"/ccpa",
|
||||||
|
"/sitemap",
|
||||||
|
"/archive",
|
||||||
|
"/tags",
|
||||||
|
"/categories",
|
||||||
|
"/search",
|
||||||
|
"/users",
|
||||||
|
"/accounts",
|
||||||
|
"/dashboard",
|
||||||
|
"/profile",
|
||||||
|
"/settings",
|
||||||
|
"/preferences",
|
||||||
|
"/notifications",
|
||||||
|
"/messages",
|
||||||
|
"/inbox",
|
||||||
|
"/calendar",
|
||||||
|
"/reports",
|
||||||
|
"/analytics",
|
||||||
|
"/billing",
|
||||||
|
"/invoice",
|
||||||
|
"/orders",
|
||||||
|
"/cart",
|
||||||
|
"/checkout",
|
||||||
|
"/store",
|
||||||
|
"/shop",
|
||||||
|
"/home",
|
||||||
|
"/main",
|
||||||
|
"/start",
|
||||||
|
"/welcome",
|
||||||
|
"/intro",
|
||||||
|
"/overview",
|
||||||
|
"/summary",
|
||||||
|
"/portfolio",
|
||||||
|
"/projects",
|
||||||
|
"/work",
|
||||||
|
"/solutions",
|
||||||
|
]
|
||||||
|
|
||||||
|
ABUSE_USER_AGENTS: list[str] = [
|
||||||
|
# Desktop browsers
|
||||||
|
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
|
||||||
|
"(KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
|
||||||
|
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 "
|
||||||
|
"(KHTML, like Gecko) Version/17.5 Safari/605.1.15",
|
||||||
|
"Mozilla/5.0 (X11; Linux x86_64; rv:130.0) Gecko/20100101 Firefox/130.0",
|
||||||
|
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:130.0) Gecko/20100101 Firefox/130.0",
|
||||||
|
"Mozilla/5.0 (Linux; Android 14; SM-S918B) AppleWebKit/537.36 "
|
||||||
|
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36",
|
||||||
|
"Mozilla/5.0 (iPhone; CPU iPhone OS 17_5 like Mac OS X) AppleWebKit/605.1.15 "
|
||||||
|
"(KHTML, like Gecko) Version/17.5 Mobile/15E148 Safari/604.1",
|
||||||
|
# Well-known crawlers / bots
|
||||||
|
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; "
|
||||||
|
"+http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
|
||||||
|
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; "
|
||||||
|
"+http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
|
||||||
|
"Mozilla/5.0 (compatible; DuckDuckBot/1.1; +http://duckduckgo.com/duckduckbot.html)",
|
||||||
|
"Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)",
|
||||||
|
"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 "
|
||||||
|
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36 "
|
||||||
|
"(compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
|
||||||
|
"Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)",
|
||||||
|
"Mozilla/5.0 (compatible; DotBot/1.2; +https://opensiteexplorer.org/dotbot; help@moz.com)",
|
||||||
|
"Mozilla/5.0 (compatible; SemrushBot/7~bl; +http://www.semrush.com/bot.html)",
|
||||||
|
"Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)",
|
||||||
|
# Social / service fetchers
|
||||||
|
"facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)",
|
||||||
|
"Twitterbot/1.0",
|
||||||
|
"LinkedInBot/1.0 (compatible; Mozilla/5.0; Apache-HttpClient +http://www.linkedin.com)",
|
||||||
|
"Slackbot-LinkExpanding 1.0 (+https://api.slack.com/robots)",
|
||||||
|
"WhatsApp/2.23.20.0",
|
||||||
|
# Command-line / library clients
|
||||||
|
"curl/8.5.0",
|
||||||
|
"Wget/1.21.4 (linux-gnu)",
|
||||||
|
"python-requests/2.32.3",
|
||||||
|
"Go-http-client/1.1",
|
||||||
|
"Node.js/20.5.1",
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def _random_ipv6_host(prefix: str) -> str:
|
||||||
|
"""Return a concrete address within an IPv6 /64 prefix.
|
||||||
|
|
||||||
|
The host part is generated randomly, mimicking a fresh OS privacy address.
|
||||||
|
The input prefix must end in ``::/64`` (e.g. ``2a02:8109:ac82:6f0c::/64``).
|
||||||
|
"""
|
||||||
|
if "/" not in prefix:
|
||||||
|
return prefix
|
||||||
|
base, mask = prefix.split("/")
|
||||||
|
if mask != "64":
|
||||||
|
raise ValueError(f"only /64 IPv6 prefixes are supported, got {prefix!r}")
|
||||||
|
if base.endswith("::"):
|
||||||
|
base = base[:-2]
|
||||||
|
host = ":".join(f"{random.randint(0, 0xffff):04x}" for _ in range(4))
|
||||||
|
return f"{base}:{host}"
|
||||||
|
|
||||||
|
|
||||||
|
def _concretize_ip(entry: str) -> str:
|
||||||
|
"""Return a concrete IP address; randomise the host part for IPv6 /64 prefixes."""
|
||||||
|
if ":" in entry and "/" in entry:
|
||||||
|
return _random_ipv6_host(entry)
|
||||||
|
return entry
|
||||||
|
|
||||||
|
|
||||||
|
class _SessionIP:
|
||||||
|
"""Stable IP for a browser session, with one optional mid-session rotation.
|
||||||
|
|
||||||
|
IPv6 prefixes get a fresh random host part; IPv4 addresses are swapped for
|
||||||
|
another address from the residential pool.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, entry: str, pool: Sequence[str]):
|
||||||
|
self.entry = entry
|
||||||
|
self.pool = pool
|
||||||
|
self._value = _concretize_ip(entry)
|
||||||
|
|
||||||
|
def current(self) -> str:
|
||||||
|
return self._value
|
||||||
|
|
||||||
|
def rotate(self) -> None:
|
||||||
|
if ":" in self.entry and "/" in self.entry:
|
||||||
|
self._value = _random_ipv6_host(self.entry)
|
||||||
|
return
|
||||||
|
# IPv4: switch to another IPv4 address from the residential pool.
|
||||||
|
for _ in range(20):
|
||||||
|
candidate_entry = random.choice(self.pool)
|
||||||
|
if ":" in candidate_entry and "/" in candidate_entry:
|
||||||
|
continue
|
||||||
|
candidate = _concretize_ip(candidate_entry)
|
||||||
|
if candidate != self._value:
|
||||||
|
self._value = candidate
|
||||||
|
return
|
||||||
|
|
||||||
|
|
||||||
def _sleep(base: float, jitter: float) -> None:
|
def _sleep(base: float, jitter: float) -> None:
|
||||||
time.sleep(max(0.0, base + random.uniform(-jitter, jitter)))
|
time.sleep(max(0.0, base + random.uniform(-jitter, jitter)))
|
||||||
|
|
||||||
|
|
||||||
def _source_ip(index: int) -> str:
|
def _normalize_url(url: str) -> str:
|
||||||
"""Pick one of the small pool of real public IPs."""
|
"""Return a usable base URL, adding missing scheme/host/port parts.
|
||||||
return SOURCE_IPS[index % len(SOURCE_IPS)]
|
|
||||||
|
- bare ``:PORT`` becomes ``http://localhost:PORT``
|
||||||
|
- missing scheme becomes ``http://``
|
||||||
|
- otherwise returned as-is
|
||||||
|
|
||||||
|
Raises ``ValueError`` when the result is not a valid http(s) URL.
|
||||||
|
"""
|
||||||
|
raw = url.strip()
|
||||||
|
if not raw:
|
||||||
|
raise ValueError("empty URL")
|
||||||
|
if raw.startswith(":"):
|
||||||
|
raw = f"http://localhost{raw}"
|
||||||
|
elif raw.isdigit():
|
||||||
|
raw = f"http://localhost:{raw}"
|
||||||
|
elif not raw.startswith(("http://", "https://")):
|
||||||
|
raw = f"http://{raw}"
|
||||||
|
parsed = urlparse(raw)
|
||||||
|
if parsed.scheme not in ("http", "https") or not parsed.netloc:
|
||||||
|
raise ValueError(f"invalid URL: {url!r}")
|
||||||
|
return raw
|
||||||
|
|
||||||
|
|
||||||
def _poisson_wait(rate: float) -> float:
|
def _poisson_wait(rate: float) -> float:
|
||||||
@@ -132,31 +513,39 @@ def _poisson_wait(rate: float) -> float:
|
|||||||
return random.expovariate(rate)
|
return random.expovariate(rate)
|
||||||
|
|
||||||
|
|
||||||
def _collect_links(page: Any) -> list[dict[str, Any]]:
|
def _collect_links(page: Any, include_external: bool = False) -> list[dict[str, Any]]:
|
||||||
"""Return internal links from the current page, excluding the current page."""
|
"""Return links from the current page, excluding the current page.
|
||||||
|
|
||||||
|
Internal links stay on the site; external links are real https URLs found
|
||||||
|
in the page content and are marked with ``external: true``.
|
||||||
|
"""
|
||||||
return page.evaluate(
|
return page.evaluate(
|
||||||
"""() => {
|
"""(includeExternal) => {
|
||||||
const loc = new URL(location.href);
|
const loc = new URL(location.href);
|
||||||
return Array.from(document.querySelectorAll('a[href]'))
|
const out = [];
|
||||||
.filter(a => {
|
for (const a of document.querySelectorAll('a[href]')) {
|
||||||
try {
|
try {
|
||||||
const u = new URL(a.href);
|
const u = new URL(a.href);
|
||||||
return u.origin === loc.origin
|
|
||||||
&& !u.pathname.startsWith('/_')
|
|
||||||
&& !u.pathname.startsWith('/auth')
|
|
||||||
&& u.pathname !== '/favicon.ico'
|
|
||||||
&& u.pathname !== loc.pathname;
|
|
||||||
} catch { return false; }
|
|
||||||
})
|
|
||||||
.map(a => {
|
|
||||||
const rect = a.getBoundingClientRect();
|
const rect = a.getBoundingClientRect();
|
||||||
return {
|
const item = {
|
||||||
href: a.href,
|
href: a.href,
|
||||||
text: (a.innerText || a.title || '').trim().slice(0, 60),
|
text: (a.innerText || a.title || '').trim().slice(0, 60),
|
||||||
visible: !!(rect.width && rect.height && rect.top < window.innerHeight && rect.bottom > 0),
|
visible: !!(rect.width && rect.height && rect.top < window.innerHeight && rect.bottom > 0),
|
||||||
};
|
};
|
||||||
});
|
if (u.origin === loc.origin
|
||||||
}"""
|
&& !u.pathname.startsWith('/_')
|
||||||
|
&& !u.pathname.startsWith('/auth')
|
||||||
|
&& u.pathname !== '/favicon.ico'
|
||||||
|
&& u.pathname !== loc.pathname) {
|
||||||
|
out.push(item);
|
||||||
|
} else if (includeExternal && u.protocol === 'https:' && u.origin !== loc.origin) {
|
||||||
|
out.push({ ...item, external: true });
|
||||||
|
}
|
||||||
|
} catch { /* ignore malformed hrefs */ }
|
||||||
|
}
|
||||||
|
return out;
|
||||||
|
}""",
|
||||||
|
include_external,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
@@ -197,37 +586,75 @@ def _run_browser_session(
|
|||||||
paths: Sequence[str],
|
paths: Sequence[str],
|
||||||
profile: BrowserProfile,
|
profile: BrowserProfile,
|
||||||
session_index: int,
|
session_index: int,
|
||||||
max_clicks: int,
|
ip_entry: str,
|
||||||
stay: tuple[float, float],
|
|
||||||
headless: bool,
|
|
||||||
fake_ip: str,
|
|
||||||
) -> dict[str, Any]:
|
) -> dict[str, Any]:
|
||||||
from playwright.sync_api import sync_playwright
|
from playwright.sync_api import sync_playwright
|
||||||
|
|
||||||
|
MAX_CLICKS = 6
|
||||||
|
STAY = (2.0, 6.0)
|
||||||
|
HEADLESS = True
|
||||||
|
REFERER_RATE = 0.75
|
||||||
|
INCLUDE_EXTERNAL = True
|
||||||
|
|
||||||
|
ip_provider = _SessionIP(ip_entry, RESIDENTIAL_SOURCE_IPS)
|
||||||
|
ips_used: list[str] = [ip_provider.current()]
|
||||||
trail: list[str] = []
|
trail: list[str] = []
|
||||||
start_time = datetime.now(UTC)
|
start_time = datetime.now(UTC)
|
||||||
try:
|
try:
|
||||||
with sync_playwright() as p:
|
with sync_playwright() as p:
|
||||||
browser = p.chromium.launch(
|
browser = p.chromium.launch(
|
||||||
headless=headless,
|
headless=HEADLESS,
|
||||||
args=["--no-sandbox", "--disable-dev-shm-usage"],
|
args=["--no-sandbox", "--disable-dev-shm-usage"],
|
||||||
)
|
)
|
||||||
|
extra_headers = {
|
||||||
|
"X-Forwarded-For": ip_provider.current(),
|
||||||
|
"Accept-Language": profile.accept_language,
|
||||||
|
}
|
||||||
|
# Most sessions arrive from an external origin; some are direct.
|
||||||
|
# A subset of referered sessions carries realistic UTM tags on the
|
||||||
|
# landing URL; the referer origin is paired with the UTM source.
|
||||||
|
tagged: dict[str, str] = {}
|
||||||
|
if random.random() < REFERER_RATE:
|
||||||
|
if random.random() < UTM_RATE:
|
||||||
|
referer, tagged = random.choice(TAGGED_REFERRERS)
|
||||||
|
else:
|
||||||
|
referer = random.choice(PLAIN_REFERRERS)
|
||||||
|
extra_headers["Referer"] = referer
|
||||||
context = browser.new_context(
|
context = browser.new_context(
|
||||||
user_agent=profile.user_agent,
|
user_agent=profile.user_agent,
|
||||||
viewport={"width": profile.viewport[0], "height": profile.viewport[1]},
|
viewport={"width": profile.viewport[0], "height": profile.viewport[1]},
|
||||||
extra_http_headers={
|
extra_http_headers=extra_headers,
|
||||||
"X-Forwarded-For": fake_ip,
|
|
||||||
"Accept-Language": profile.accept_language,
|
|
||||||
},
|
|
||||||
)
|
)
|
||||||
page = context.new_page()
|
page = context.new_page()
|
||||||
|
|
||||||
|
# Update X-Forwarded-For per request; the value stays stable unless we
|
||||||
|
# explicitly rotate it once mid-session.
|
||||||
|
def _route_handler(route, request):
|
||||||
|
headers = dict(request.headers)
|
||||||
|
headers["X-Forwarded-For"] = ip_provider.current()
|
||||||
|
ips_used.append(headers["X-Forwarded-For"])
|
||||||
|
route.continue_(headers=headers)
|
||||||
|
|
||||||
|
page.route("**/*", _route_handler)
|
||||||
|
|
||||||
|
# Pick one point during the session to emulate an IP rotation.
|
||||||
|
rotate_at = random.randint(0, MAX_CLICKS - 1) if MAX_CLICKS > 0 else -1
|
||||||
|
|
||||||
entry = random.choice(paths) if paths else "/"
|
entry = random.choice(paths) if paths else "/"
|
||||||
page.goto(urljoin(base, entry), wait_until="networkidle")
|
landing = urljoin(base, entry)
|
||||||
|
if tagged:
|
||||||
|
sep = "&" if "?" in landing else "?"
|
||||||
|
landing += sep + urlencode(tagged)
|
||||||
|
page.goto(landing, wait_until="networkidle")
|
||||||
trail.append(page.url)
|
trail.append(page.url)
|
||||||
|
|
||||||
for _ in range(max_clicks):
|
for click_idx in range(MAX_CLICKS):
|
||||||
_sleep(random.uniform(*stay) / 2, 0.3)
|
_sleep(random.uniform(*STAY) / 2, 0.3)
|
||||||
links = _collect_links(page)
|
if click_idx == rotate_at:
|
||||||
|
ip_provider.rotate()
|
||||||
|
ips_used.append(ip_provider.current())
|
||||||
|
logger.debug("rotated session IP to %s", ip_provider.current())
|
||||||
|
links = _collect_links(page, INCLUDE_EXTERNAL)
|
||||||
visible = [item for item in links if item.get("visible")]
|
visible = [item for item in links if item.get("visible")]
|
||||||
if not visible:
|
if not visible:
|
||||||
visible = links
|
visible = links
|
||||||
@@ -242,16 +669,23 @@ def _run_browser_session(
|
|||||||
ok = _click_link(page, alt)
|
ok = _click_link(page, alt)
|
||||||
if not ok:
|
if not ok:
|
||||||
break
|
break
|
||||||
|
if link.get("external"):
|
||||||
|
# Outbound navigation: the analytics exit ping is already
|
||||||
|
# in flight. Record the external URL and end the session.
|
||||||
|
trail.append(page.url)
|
||||||
|
_sleep(0.5, 0.2)
|
||||||
|
break
|
||||||
page.wait_for_load_state("networkidle")
|
page.wait_for_load_state("networkidle")
|
||||||
trail.append(page.url)
|
trail.append(page.url)
|
||||||
_sleep(random.uniform(*stay), 0.5)
|
_sleep(random.uniform(*STAY), 0.5)
|
||||||
|
|
||||||
browser.close()
|
browser.close()
|
||||||
|
|
||||||
return {
|
return {
|
||||||
"profile": profile.name,
|
"profile": profile.name,
|
||||||
"entry": entry,
|
"entry": entry,
|
||||||
"ip": fake_ip,
|
"ip": ips_used[0],
|
||||||
|
"ips_seen": len(set(ips_used)),
|
||||||
"pages": len(trail),
|
"pages": len(trail),
|
||||||
"trail": [urlparse(u).path or "/" for u in trail],
|
"trail": [urlparse(u).path or "/" for u in trail],
|
||||||
"duration": (datetime.now(UTC) - start_time).total_seconds(),
|
"duration": (datetime.now(UTC) - start_time).total_seconds(),
|
||||||
@@ -265,12 +699,10 @@ def _run_crawler_hit(
|
|||||||
base: str,
|
base: str,
|
||||||
paths: Sequence[str],
|
paths: Sequence[str],
|
||||||
profile: CrawlerProfile,
|
profile: CrawlerProfile,
|
||||||
profile_index: int,
|
|
||||||
session_index: int,
|
|
||||||
) -> dict[str, Any]:
|
) -> dict[str, Any]:
|
||||||
path = random.choice(paths) if paths else "/"
|
path = random.choice(paths) if paths else "/"
|
||||||
url = urljoin(base, path)
|
url = urljoin(base, path)
|
||||||
fake_ip = _source_ip(session_index)
|
fake_ip = profile.ip
|
||||||
headers = {
|
headers = {
|
||||||
"User-Agent": profile.user_agent,
|
"User-Agent": profile.user_agent,
|
||||||
"X-Forwarded-For": fake_ip,
|
"X-Forwarded-For": fake_ip,
|
||||||
@@ -290,49 +722,99 @@ def _run_crawler_hit(
|
|||||||
return {"profile": profile.name, "path": path, "error": str(exc)}
|
return {"profile": profile.name, "path": path, "error": str(exc)}
|
||||||
|
|
||||||
|
|
||||||
|
def _abuse_ua() -> str:
|
||||||
|
"""Return a randomized, syntactically valid user agent for an abuse scan."""
|
||||||
|
return random.choice(ABUSE_USER_AGENTS)
|
||||||
|
|
||||||
|
|
||||||
|
def _run_abuse_scanner(base: str, ip_index: int) -> dict[str, Any]:
|
||||||
|
"""Fire a burst of vulnerability probes from a single fake IP.
|
||||||
|
|
||||||
|
Scanner 0 randomises its user agent every request, scanner 1 uses a fixed
|
||||||
|
browser UA, and scanner 2 uses a fixed crawler UA.
|
||||||
|
"""
|
||||||
|
ip_entry = ABUSE_SOURCE_IPS[ip_index % len(ABUSE_SOURCE_IPS)]
|
||||||
|
if ":" in ip_entry and "/" in ip_entry:
|
||||||
|
fake_ip = _random_ipv6_host(ip_entry)
|
||||||
|
else:
|
||||||
|
fake_ip = ip_entry
|
||||||
|
|
||||||
|
MIN_HITS = 15
|
||||||
|
MAX_HITS = 25
|
||||||
|
total_hits = random.randint(MIN_HITS, MAX_HITS)
|
||||||
|
|
||||||
|
# Ensure the burst contains both telltales: suspicious paths and more
|
||||||
|
# than ten normal-looking 404 paths.
|
||||||
|
suspicious_count = max(5, total_hits // 3)
|
||||||
|
normal_count = total_hits - suspicious_count
|
||||||
|
if normal_count < 11:
|
||||||
|
normal_count = 11
|
||||||
|
suspicious_count = max(3, total_hits - normal_count)
|
||||||
|
|
||||||
|
paths = random.choices(SUSPICIOUS_PATHS, k=suspicious_count) + random.choices(
|
||||||
|
NORMAL_404_PATHS, k=normal_count
|
||||||
|
)
|
||||||
|
random.shuffle(paths)
|
||||||
|
|
||||||
|
ua_mode = ip_index % 3
|
||||||
|
if ua_mode == 0:
|
||||||
|
get_ua = _abuse_ua
|
||||||
|
elif ua_mode == 1:
|
||||||
|
def get_ua() -> str:
|
||||||
|
return BROWSER_PROFILES[0].user_agent
|
||||||
|
else:
|
||||||
|
def get_ua() -> str:
|
||||||
|
return CRAWLER_PROFILES[0].user_agent
|
||||||
|
|
||||||
|
scan_results: list[dict[str, Any]] = []
|
||||||
|
with httpx.Client(follow_redirects=True, timeout=15.0) as client:
|
||||||
|
for path in paths:
|
||||||
|
headers = {
|
||||||
|
"User-Agent": get_ua(),
|
||||||
|
"X-Forwarded-For": fake_ip,
|
||||||
|
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
|
||||||
|
"Accept-Language": random.choice(
|
||||||
|
["en-US,en;q=0.9", "en-GB,en;q=0.8", "en;q=0.7"]
|
||||||
|
),
|
||||||
|
}
|
||||||
|
try:
|
||||||
|
r = client.get(urljoin(base, path), headers=headers)
|
||||||
|
scan_results.append(
|
||||||
|
{"path": path, "status": r.status_code, "ua": headers["User-Agent"]}
|
||||||
|
)
|
||||||
|
except Exception as exc: # noqa: BLE001
|
||||||
|
scan_results.append({"path": path, "error": str(exc)})
|
||||||
|
_sleep(0.15, 0.1)
|
||||||
|
|
||||||
|
return {
|
||||||
|
"scanner": ip_index + 1,
|
||||||
|
"ip": fake_ip,
|
||||||
|
"hits": len(scan_results),
|
||||||
|
"results": scan_results,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
|
def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
|
||||||
parser = argparse.ArgumentParser(
|
parser = argparse.ArgumentParser(
|
||||||
description="Generate fake traffic for a Pagerite site.",
|
description="Generate fake traffic for a Pagerite site.",
|
||||||
formatter_class=argparse.ArgumentDefaultsHelpFormatter,
|
formatter_class=argparse.ArgumentDefaultsHelpFormatter,
|
||||||
)
|
)
|
||||||
parser.add_argument("url", help="Base URL of the Pagerite site")
|
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"-b",
|
"url",
|
||||||
"--browsers",
|
nargs="?",
|
||||||
type=int,
|
default="http://localhost:8200",
|
||||||
default=5,
|
help="Base URL of the Pagerite site (default: http://localhost:8200). "
|
||||||
help="Number of simulated browser sessions",
|
"A bare :PORT or PORT is treated as http://localhost:PORT; a "
|
||||||
|
"missing scheme defaults to http://.",
|
||||||
)
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"-c", "--crawlers", type=int, default=10, help="Number of crawler HTTP GETs"
|
"-t",
|
||||||
)
|
"--duration",
|
||||||
parser.add_argument(
|
|
||||||
"--max-clicks",
|
|
||||||
type=int,
|
|
||||||
default=6,
|
|
||||||
help="Max internal link clicks per browser session",
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--stay",
|
|
||||||
type=float,
|
type=float,
|
||||||
nargs=2,
|
default=60.0,
|
||||||
default=[2.0, 6.0],
|
metavar="SECONDS",
|
||||||
metavar=("MIN", "MAX"),
|
help="Rough maximum time to generate traffic (0 runs one preset batch)",
|
||||||
help="Seconds to stay on a page before clicking again",
|
|
||||||
)
|
)
|
||||||
parser.add_argument(
|
|
||||||
"--headless",
|
|
||||||
action=argparse.BooleanOptionalAction,
|
|
||||||
default=True,
|
|
||||||
help="Run browsers headlessly",
|
|
||||||
)
|
|
||||||
parser.add_argument(
|
|
||||||
"--arrival-rate",
|
|
||||||
type=float,
|
|
||||||
default=1.0,
|
|
||||||
help="Average arrivals per second (Poisson). 0 disables inter-arrival waits",
|
|
||||||
)
|
|
||||||
parser.add_argument("--seed", type=int, default=None, help="Random seed")
|
|
||||||
parser.add_argument("-v", "--verbose", action="store_true", help="Debug logging")
|
parser.add_argument("-v", "--verbose", action="store_true", help="Debug logging")
|
||||||
return parser.parse_args(argv)
|
return parser.parse_args(argv)
|
||||||
|
|
||||||
@@ -342,8 +824,11 @@ def main(argv: Sequence[str] | None = None) -> int:
|
|||||||
if args.verbose:
|
if args.verbose:
|
||||||
logger.setLevel(logging.DEBUG)
|
logger.setLevel(logging.DEBUG)
|
||||||
|
|
||||||
random.seed(args.seed)
|
try:
|
||||||
base = args.url.rstrip("/")
|
base = _normalize_url(args.url).rstrip("/")
|
||||||
|
except ValueError as exc:
|
||||||
|
logger.error("%s", exc)
|
||||||
|
return 2
|
||||||
|
|
||||||
# Discover content paths from the public page tree if we can.
|
# Discover content paths from the public page tree if we can.
|
||||||
paths: list[str] = []
|
paths: list[str] = []
|
||||||
@@ -357,59 +842,95 @@ def main(argv: Sequence[str] | None = None) -> int:
|
|||||||
paths = ["/"]
|
paths = ["/"]
|
||||||
|
|
||||||
logger.info(
|
logger.info(
|
||||||
"Generating fake traffic against %s (%d content paths, %d browsers, %d crawlers)",
|
"Generating fake traffic against %s (%d content paths, duration=%ss)",
|
||||||
base,
|
base,
|
||||||
len(paths),
|
len(paths),
|
||||||
args.browsers,
|
args.duration,
|
||||||
args.crawlers,
|
|
||||||
)
|
)
|
||||||
|
|
||||||
results: list[dict[str, Any]] = []
|
results: list[dict[str, Any]] = []
|
||||||
|
arrival_rate = 1.0
|
||||||
|
|
||||||
for i in range(args.browsers):
|
def _wait() -> None:
|
||||||
if i > 0:
|
wait = _poisson_wait(arrival_rate)
|
||||||
wait = _poisson_wait(args.arrival_rate)
|
logger.debug("waiting %.2fs before next session", wait)
|
||||||
logger.debug("waiting %.2fs before next browser session", wait)
|
time.sleep(wait)
|
||||||
time.sleep(wait)
|
|
||||||
profile = random.choice(BROWSER_PROFILES)
|
|
||||||
fake_ip = _source_ip(i)
|
|
||||||
logger.info(
|
|
||||||
"[%d/%d] browser session: %s (ip=%s)",
|
|
||||||
i + 1,
|
|
||||||
args.browsers,
|
|
||||||
profile.name,
|
|
||||||
fake_ip,
|
|
||||||
)
|
|
||||||
result = _run_browser_session(
|
|
||||||
base,
|
|
||||||
paths,
|
|
||||||
profile,
|
|
||||||
i,
|
|
||||||
args.max_clicks,
|
|
||||||
(args.stay[0], args.stay[1]),
|
|
||||||
args.headless,
|
|
||||||
fake_ip,
|
|
||||||
)
|
|
||||||
results.append(result)
|
|
||||||
logger.debug(" trail: %s", result.get("trail", []))
|
|
||||||
|
|
||||||
for i in range(args.crawlers):
|
if args.duration <= 0:
|
||||||
if i > 0:
|
# One preset batch.
|
||||||
wait = _poisson_wait(args.arrival_rate)
|
for i in range(5):
|
||||||
logger.debug("waiting %.2fs before next crawler hit", wait)
|
if i > 0:
|
||||||
time.sleep(wait)
|
_wait()
|
||||||
profile_index = i % len(CRAWLER_PROFILES)
|
profile = random.choice(BROWSER_PROFILES)
|
||||||
profile = CRAWLER_PROFILES[profile_index]
|
ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
|
||||||
fake_ip = _source_ip(i)
|
logger.info(
|
||||||
logger.info(
|
"browser session: %s (ip=%s)",
|
||||||
"[%d/%d] crawler hit: %s (ip=%s)",
|
profile.name,
|
||||||
i + 1,
|
_concretize_ip(ip_entry),
|
||||||
args.crawlers,
|
)
|
||||||
profile.name,
|
result = _run_browser_session(base, paths, profile, i, ip_entry)
|
||||||
fake_ip,
|
results.append(result)
|
||||||
)
|
logger.debug(" trail: %s", result.get("trail", []))
|
||||||
result = _run_crawler_hit(base, paths, profile, profile_index, i)
|
|
||||||
results.append(result)
|
for i in range(10):
|
||||||
|
if i > 0:
|
||||||
|
_wait()
|
||||||
|
profile = random.choice(CRAWLER_PROFILES)
|
||||||
|
logger.info(
|
||||||
|
"crawler hit: %s (ip=%s)",
|
||||||
|
profile.name,
|
||||||
|
profile.ip,
|
||||||
|
)
|
||||||
|
result = _run_crawler_hit(base, paths, profile)
|
||||||
|
results.append(result)
|
||||||
|
|
||||||
|
for i in range(3):
|
||||||
|
if i > 0:
|
||||||
|
_wait()
|
||||||
|
ip_entry = ABUSE_SOURCE_IPS[i % len(ABUSE_SOURCE_IPS)]
|
||||||
|
logger.info("abuse scanner: %s", ip_entry)
|
||||||
|
result = _run_abuse_scanner(base, i)
|
||||||
|
results.append(result)
|
||||||
|
logger.debug(
|
||||||
|
" hits: %s", [r.get("path") for r in result.get("results", [])]
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
deadline = time.time() + args.duration
|
||||||
|
session_index = 0
|
||||||
|
while time.time() < deadline:
|
||||||
|
if session_index > 0:
|
||||||
|
_wait()
|
||||||
|
phase = session_index % 3
|
||||||
|
if phase == 0:
|
||||||
|
profile = random.choice(BROWSER_PROFILES)
|
||||||
|
ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
|
||||||
|
logger.info(
|
||||||
|
"browser session: %s (ip=%s)",
|
||||||
|
profile.name,
|
||||||
|
_concretize_ip(ip_entry),
|
||||||
|
)
|
||||||
|
result = _run_browser_session(
|
||||||
|
base, paths, profile, session_index, ip_entry
|
||||||
|
)
|
||||||
|
logger.debug(" trail: %s", result.get("trail", []))
|
||||||
|
elif phase == 1:
|
||||||
|
profile = random.choice(CRAWLER_PROFILES)
|
||||||
|
logger.info(
|
||||||
|
"crawler hit: %s (ip=%s)",
|
||||||
|
profile.name,
|
||||||
|
profile.ip,
|
||||||
|
)
|
||||||
|
result = _run_crawler_hit(base, paths, profile)
|
||||||
|
else:
|
||||||
|
ip_entry = ABUSE_SOURCE_IPS[session_index % len(ABUSE_SOURCE_IPS)]
|
||||||
|
logger.info("abuse scanner: %s", ip_entry)
|
||||||
|
result = _run_abuse_scanner(base, session_index // 3)
|
||||||
|
logger.debug(
|
||||||
|
" hits: %s",
|
||||||
|
[r.get("path") for r in result.get("results", [])],
|
||||||
|
)
|
||||||
|
results.append(result)
|
||||||
|
session_index += 1
|
||||||
|
|
||||||
ok = sum(1 for r in results if "error" not in r)
|
ok = sum(1 for r in results if "error" not in r)
|
||||||
logger.info("Done: %d/%d requests succeeded.", ok, len(results))
|
logger.info("Done: %d/%d requests succeeded.", ok, len(results))
|
||||||
|
|||||||
Reference in New Issue
Block a user