Compare commits

..
17 Commits
Author SHA1 Message Date
LeoVasanko 29f8fac013 Slightly prettier analytics URL 2026-08-21 23:58:00 +00:00
LeoVasanko 6199e5a69e Support for UTM tags in transition graph as source sites. 2026-08-21 23:50:01 +00:00
LeoVasanko 375b4b6bdb analytics: unify visitor cell across visits, crawlers and abuse tables 2026-08-21 23:32:46 +00:00
LeoVasanko fdb3e42d6f analytics: shared Client struct, grouped abuse paths, unified visitor cell 2026-08-21 23:16:15 +00:00
LeoVasanko 51a6a16221 Neater abuse table formatting. 2026-08-21 22:31:37 +00:00
LeoVasanko 0798e24d24 Desaturated house emojis 2026-08-21 22:03:55 +00:00
LeoVasanko 3be2d08ac9 SI formatting of large visitor numbers. 2026-08-21 21:36:57 +00:00
LeoVasanko b7d5b23ae6 Cleaner formatting of utm tags in visitor table. 2026-08-21 21:22:16 +00:00
LeoVasanko 6eaa1c1a8b analytics: 24h day view with bar chart for precise realtime stats. Tables redesigned with cleaner layout. Tracking article read times. Adjust connection graph visualizations by time range. Other cleanup and supporting systems. 2026-08-21 20:14:04 +00:00
LeoVasanko f341d22aa0 Improved fake traffic generation with abuse bots, utm tags etc. 2026-08-21 20:10:41 +00:00
LeoVasanko 1a479ceb24 Add --dbip CLI flag to auto-download/update the DB-IP MMDB database.
Downloads the latest dbip-city-lite-YYYY-MM.mmdb.gz before starting the
server, skipping when the local database is current, falling back to the
previous month on 404, and removing older databases after an update.
Promotes httpx to a runtime dependency.
2026-08-21 03:09:04 +00:00
LeoVasanko ff553d018a Default scheme, host and port for fake_traffic script. 2026-08-21 02:52:53 +00:00
LeoVasanko c807d48a13 Add more external content in seed data. 2026-08-21 02:50:36 +00:00
LeoVasanko 9c383c1c8b Change default port mapping to 8100/8200/8210 (prod/vite/dev). Vite gets different port to avoid caching problems when switching between it and prod. 2026-08-21 02:49:18 +00:00
LeoVasanko 462e995adc Add external link (referer/outgoing) display on connection graph. 2026-08-21 02:44:54 +00:00
LeoVasanko ea069b98da Fix analytics app not mounting on fetch-navigation to /_a
load() queried the live document for the pagerite:analytics-src meta,
but the swap never touches <head> — the meta only exists in the fetched
doc, so the app never mounted unless /_a was loaded directly. Also cache
the fetched HTML so the post-swap preload doesn't re-GET the page we
just navigated to.
2026-08-21 01:53:39 +00:00
LeoVasanko deb5419c47 analytics improvements:
- keep visitor charts y-axis minimum range at 10
- keep 'all' chart x-axis minimum span at 30 days
- group crawler hits by (ip, ua) and list top pages visited, show crawler page load counts as N× prefix
- store and display geoip city, keep geoip country overwrite
- stream live updates over WebSocket /_api/ws/analytics
- include family ring arcs in transition map crop bounds
- remove top UA summary, limit crawlers to 10 and visits to 20
- human-readable relative timestamps with UTC tooltip
2026-08-21 01:36:58 +00:00
20 changed files with 2881 additions and 662 deletions
+129 -62
View File
@@ -5,11 +5,12 @@ Struct dumped to disk — separate from the kanta content database, path from
`PAGERITE_ANALYTICS` (default: the database path with `.kantadb` replaced by `PAGERITE_ANALYTICS` (default: the database path with `.kantadb` replaced by
`.analytics.json`, e.g. `pagerite.analytics.json`). `.analytics.json`, e.g. `pagerite.analytics.json`).
- `pagerite/analytics.py` — data model (`Analytics`, `Visit`) and the `Store` - `pagerite/analytics.py` — data model (`Analytics`, `Client`, `Visit`,
(in-memory data + session map, atomic JSON persistence). `CrawlerHit`, `AbuseHit`) and the `Store` (in-memory data + session map,
atomic JSON persistence).
- `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`), - `pagerite/app.py` — entry-referer stashing in `show_page` (`_track_entry`),
the `POST /_a` ping endpoint, and `GET /_api/analytics` (admin-gated like the `POST /_a` ping endpoint, and `WebSocket /_api/ws/analytics`
every `/_api` endpoint). (admin-gated like every `/_api` endpoint).
- `frontend/src/pagerite.js` — client navigation pings and the 📊 pen. - `frontend/src/pagerite.js` — client navigation pings and the 📊 pen.
- `frontend/src/AnalyticsView.vue` — viewer component rendered inside the - `frontend/src/AnalyticsView.vue` — viewer component rendered inside the
normal site layout on the `/_a` analytics page. normal site layout on the `/_a` analytics page.
@@ -32,76 +33,133 @@ The client (`pagerite.js`) POSTs fire-and-forget pings to `/_a` with
- **Internal fetch-navigations**: `to` is the target path, sent only after - **Internal fetch-navigations**: `to` is the target path, sent only after
the swap actually happened (a failed swap falls back to a full load, the swap actually happened (a failed swap falls back to a full load,
whose initial ping counts the view instead — no gap, no double count). whose initial ping counts the view instead — no gap, no double count).
- **External links** (`https` only): `to` is the link's origin. This is the - **External links** (`https` only): `to` is the link's full URL. This is the
exit-link record; the user may continue navigating afterwards (new tab, exit-link record; the user may continue navigating afterwards (new tab,
back), so the exit origin is not necessarily the last trail entry. back), so the exit URL is not necessarily the last trail entry. Outbound
links are stored by full URL so several links to the same domain remain
distinct.
- **Excluded**: back/forward (popstate) navigations, navigation involving - **Excluded**: back/forward (popstate) navigations, navigation involving
the analytics page itself (`/_a`), and everything while the user is known to the analytics page itself (`/_a`), and everything while the user has the
be an admin *and SSO is actually in use* — with no auth proxy (dev/test) editor open (`body.editing`). Admin noise, not visits.
"admin" is everyone's state, so the gate is off and everything is recorded — - **Admins**: when SSO is in use and the session is known to be an admin,
or has the editor open (`body.editing`). Admin noise, not visits. the client still pings but adds `hide=1`. The server then records
nothing — and if the same client session already had a visit from before
logging in, that visit is removed from the JSON along with the counts
recorded when it was created (site visit, entry view, entry transition).
Views/transitions logged by later pings inside such a visit lack
per-event timestamps and are left as-is. With no auth proxy (dev/test)
"admin" is everyone's state, so `hide` stays 0 and everything is recorded.
- The server validates `to`: internal paths must be valid slug paths - The server validates `to`: internal paths must be valid slug paths
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the ("/" or `[a-z0-9_-]` segments), external ones are re-derived to the
https origin and accepted only when the client sent exactly that. https origin and accepted only when the client sent exactly that.
- The initial ping also records the visitor's `User-Agent` and - **Client records**: the visitor's IP (IPv4 or IPv6 /64 network), raw
`Accept-Language` headers. The first `Accept-Language` tag is stored as `User-Agent` and extracted `Accept-Language` tag are hashed with blake3;
`lang` (e.g. `en-us`) and its region subtag, if present, is stored as the first 6 bytes identify a shared `Client` record. The `Client` stores
an initial `country` (e.g. `US`). the full IP, `User-Agent`, compact `ua_pretty`, `lang`, initial
- The visitor IP is stored. A reverse-DNS lookup is attempted for each new `country` from the language-region subtag, and asynchronously-filled
visit and the result, when available, is cached in RAM and stored as `country`/`city` from DB-IP geoip plus reverse-DNS `host`. Visits,
`host`; local/reserved/multicast addresses are skipped. crawler hits and abuse hits all reference this record by its hash, so
- If a DB-IP MMDB file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the client metadata is stored once instead of repeated per event.
repository root, it is loaded at startup and used to look up a more accurate - The visitor IP is stored in the `Client`. A reverse-DNS lookup is
`country`. The MMDB lookup and the reverse-DNS lookup run in background attempted for each new client and the result, when available, is stored as
tasks after the visit is stored, so the `/ _a` response is never delayed. `host`; local/reserved/multicast addresses are skipped. If a DB-IP MMDB
The decompressed `dbip-*.mmdb` file is kept in the repository root and file (`dbip-*.mmdb` or `dbip-*.mmdb.gz`) is present in the repository
ignored by git. root, it is loaded at startup and used to look up `country`/`city`. These
lookups run in background tasks after the event is stored, so the `/_a`
response is never delayed. The decompressed `dbip-*.mmdb` file is kept in
the repository root and ignored by git. The CLI flag `--dbip`
(`uv run pagerite --dbip`) downloads the latest
`dbip-city-lite-YYYY-MM.mmdb.gz` from DB-IP before the server starts,
skipping the download when the local database is already current and
removing older versions after an update; without the flag only an existing
file is used.
- **Crawler hits**: every document GET is queued in RAM as a pending crawler - **Crawler hits**: every document GET is queued in RAM as a pending crawler
hit. If a ping from the same (IP, User-Agent) pair arrives within 10 hit. If a ping from the same client arrives within 10 seconds the hit is
seconds the hit is discarded; otherwise it is written to `crawlers`. discarded; otherwise it is written to `crawlers`. Crawlers do not count as
Crawlers do not count as visits or views. visits or views. The `Accept-Language` header is stored on the shared
`Client` immediately; reverse-DNS host names and DB-IP geoip
country/city are filled in asynchronously, just like for real visits. In
the analytics viewer, crawler hits are grouped by client hash and shown as
a trail of internal pages that crawler visited; the crawler table lists
the most active crawlers first rather than the most recent hits.
- **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment
starting with a dot (`/.env`, `/.git/config`) or ending in `.php`
classifies the source IP as abuse immediately, and ten plain 404s from one
IP do too. Classification reclassifies history: all earlier crawler hits
from that IP (persisted and pending) move to the `abuse` list, so a
random-UA scanner no longer pollutes the crawler stats of the legitimate
bot it impersonates. Once classified, every document GET and 404 from the
IP is recorded as an abuse hit with the full request path (query string
included), and its pings are ignored. The classified IP set (`abuse_ips`)
is persisted in the JSON file; the plain-404 counters are RAM-only. In the
viewer, abuse hits are grouped by IP (never by client/UA — scanners
randomize theirs) in a separate "Abuse" table. Identical paths are
collapsed into one entry with their hit count; flagged paths that
triggered classification are lifted to the top, followed by other 404s and
then document GETs from the abuser. Raw User-Agent strings are shown one
per line with their occurrence counts, and the full lists are click-to-copy.
## Visits and sessions ## Visits and sessions
There are no cookies. A visit is tied together by the (IP, User-Agent) pair There are no cookies. A visit is tied together by a client hash — the first
(IP from the first `X-Forwarded-For` hop — we sit behind a proxy — else the 6 bytes of a blake3 digest over the prettified IP (IPv4 unchanged, IPv6
direct peer): the first ping from a pair starts a new visit, subsequent /64 network), the raw `User-Agent` string and the extracted
pings extend it. Pings arriving with no known session (server restart) `Accept-Language` tag. The first ping from a client hash starts a new
start a fresh visit from the first ping — treated as missing data rather visit; subsequent pings extend it. Pings arriving with no known session
than dropped. The (IP, UA) → visit map and the IP → entry-referer/UTM (server restart) start a fresh visit from the first ping — treated as
tables are in-memory only, but the IP and any resolvable reverse-DNS host missing data rather than dropped. The client-hash → visit map and the IP →
name are stored on the `Visit` record itself. entry-referer/UTM tables are in-memory only; client metadata is stored in
`Analytics.clients` keyed by the client hash.
Each `Client` record:
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer),
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
- `lang` — first `Accept-Language` tag, lowercased (e.g. `"en-us"`),
- `country` — two-letter country code. Initially derived from the
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
when a database is available,
- `city` — city name from the DB-IP MMDB lookup, when available,
- `ua` — raw `User-Agent` string,
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
parsable, otherwise the raw string.
Each `Visit` record: Each `Visit` record:
- `start` — timestamp of the first event, - `start` — timestamp of the first event,
- `entry` — first page (path) seen, - `entry` — first page (path) seen,
- `referer` — external https origin of the initial load, `""` for direct, - `referer` — external https origin of the initial load, `""` for direct,
- `ip` — visitor IP address (first `X-Forwarded-For` hop, or direct peer), - `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `host` — reverse-DNS host name for `ip` when resolvable, else `""`,
- `trail` — everything seen afterwards in first-seen order: page paths and - `trail` — everything seen afterwards in first-seen order: page paths and
external exit origins. Re-visiting an already seen page (incl. the entry) external exit URLs. Re-visiting an already seen page (incl. the entry)
does not append. does not append.
- `lang` — first `Accept-Language` tag, lowercased (e.g. `en-us`),
- `country` — two-letter country code. Initially derived from the
`Accept-Language` region subtag, but overwritten by the DB-IP MMDB result
when a database is available,
- `ua` — raw `User-Agent` string from the initial ping,
- `ua_pretty` — compact display form of the UA (browser/OS/device) when
parsable, otherwise the raw string,
- `utm``utm_*` query parameters from the landing URL, as a dict. - `utm``utm_*` query parameters from the landing URL, as a dict.
- `read` — active reading time per path (seconds), keyed by path.
Each `CrawlerHit` record: Each `CrawlerHit` record:
- `start` — timestamp of the document GET, - `start` — timestamp of the document GET,
- `entry` — page path requested, - `entry` — page path requested,
- `ip` — IP address, - `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `ua` — raw `User-Agent` header,
- `ua_pretty` — compact display form of the UA when parsable,
- `referer` — external https origin of the request, `""` for direct/none, - `referer` — external https origin of the request, `""` for direct/none,
- `query` — raw query string of the request. - `query` — raw query string of the request.
Crawler hits are grouped by User-Agent in the analytics viewer. Each `AbuseHit` record:
- `start` — timestamp of the request,
- `path` — full request path including the query string (e.g. `/.env?x=1`),
- `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `flag` — true for the path that triggered abuse classification (telltale
path or the 404 that crossed the threshold),
- `is_404` — true for 404 responses, false for document GETs from the
abuser.
Crawler hits are grouped by client hash in the analytics viewer; abuse hits
are grouped by IP alone (resolved from the referenced `Client`). In the
Abuse table identical paths are collapsed with their counts; flagged paths
that triggered classification are lifted to the top, followed by other 404s
and then document GETs from the abuser. Within each category paths are
sorted by count descending, then by their earliest hit.
## Aggregates ## Aggregates
@@ -131,15 +189,15 @@ The 📊 pen in the banner corner (admins only, injected by pagerite.js next to
the edit pens) links to `/_a`, the analytics page. It is a normal site page: the edit pens) links to `/_a`, the analytics page. It is a normal site page:
the standard banner, navigation and footer stay in place, and the analytics the standard banner, navigation and footer stay in place, and the analytics
content is rendered inside `#main`. The page itself is public, but the data content is rendered inside `#main`. The page itself is public, but the data
still comes from `GET /_api/analytics`, which remains admin-gated like the stream comes from `WebSocket /_api/ws/analytics`, which remains admin-gated
rest of the management API; visitors without access see the viewer with a like the rest of the management API; visitors without access see the viewer
"could not be loaded" message. with a "could not be loaded" message.
Because it is a real page, fetch-navigation handles it like any other internal Because it is a real page, fetch-navigation handles it like any other internal
link: clicking the 📊 pen (or any link to `/_a`) fetches the server-rendered link: clicking the 📊 pen (or any link to `/_a`) fetches the server-rendered
HTML, swaps the dynamic regions and mounts the Vue analytics app in place. The HTML, swaps the dynamic regions and mounts the Vue analytics app in place. The
range selector updates the URL query string (`?range=week` etc.) so links to range selector updates the URL hash (`#week` etc.) so links to a specific
a specific range can be shared. range can be shared.
`AnalyticsView.vue` is no longer a full-screen overlay; the `body.analytics-open` `AnalyticsView.vue` is no longer a full-screen overlay; the `body.analytics-open`
page-chrome hiding and `#/analytics/<range>` hash routing have been removed. page-chrome hiding and `#/analytics/<range>` hash routing have been removed.
@@ -156,16 +214,19 @@ the smoothing time scale follows the unit: the month+ sigmas are 24× the
hourly ones. The y max is derived from the smoothed curves so single-bucket hourly ones. The y max is derived from the smoothed curves so single-bucket
spikes don't blow up the scale, and raw spikes are clamped into the plot. spikes don't blow up the scale, and raw spikes are clamped into the plot.
Axes always start at 0 and end at a multiple of a 1-2-5 major step (max 5 Axes always start at 0 and end at a multiple of a 1-2-5 major step (max 5
labeled intervals, minor lines at fifths when integral; the floor is 1/h). labeled intervals, minor lines at fifths when integral; the minimum y-axis
range is 10 so tiny values such as a single visit are not stretched to a
fractional scale).
The week range is aligned to Monday 00:00 UTC and overlays up to 8 previous The week range is aligned to Monday 00:00 UTC and overlays up to 8 previous
weeks in the same accent color at decreasing opacity (the current week is weeks in the same accent color at decreasing opacity (the current week is
truncated at the current bucket, never drawing fake zeroes for the future); truncated at the current bucket, never drawing fake zeroes for the future);
its x labels are weekday names centered at midday UTC, without vertical grid its x labels are weekday names centered at midday UTC, without vertical grid
lines (day boundaries would be misleading in the viewer's timezone). The lines (day boundaries would be misleading in the viewer's timezone). The
month view labels days the same lineless way — day numbers at noon UTC, month view labels days the same lineless way — day numbers at noon UTC,
with the month name substituted for the 1st. Year and all are rolling with the month name substituted for the 1st. Year is a rolling 365-day window ending at now, re-bucketed to daily points,
windows ending at now, re-bucketed to daily points, with boundary lines at with boundary lines at months/years. All uses the full data reach, but keeps
months/years. Below the charts: a radial **transition map** (all pages from at least the past 30 days so the chart never collapses to a tiny sliver when
the site is young. Below the charts: a radial **transition map** (all pages from
`/_api/pages` — front page at the center, each slug level on its own ring, `/_api/pages` — front page at the center, each slug level on its own ring,
siblings clockwise in navigation order from the top, radial gap equal to siblings clockwise in navigation order from the top, radial gap equal to
the arc spacing — opposite transition directions joined into organic the arc spacing — opposite transition directions joined into organic
@@ -175,8 +236,14 @@ carrying less than 1% of the total traffic
pruned; beads are simulated one by one in JS (requestAnimationFrame) and pruned; beads are simulated one by one in JS (requestAnimationFrame) and
flow along each edge, emitted at a rate linearly proportional flow along each edge, emitted at a rate linearly proportional
to the directional count with no in-flight limit, opposing directions to the directional count with no in-flight limit, opposing directions
offset onto parallel lanes. External referers show as a node row above the offset onto parallel lanes. External sources show as a node row above the
map, external exits as small nodes fanned outwards from their source map: each visit is attributed to `utm_campaign`, then `utm_source`, then the
page), per-page view referer origin, then any other `utm_*` tag, so UTM-tagged visits are grouped
counts, the top transitions and the 50 most recent visit trails. Data comes from `GET /_api/analytics`, which under their campaign/source value rather than the referer domain. A UTM
returns the raw JSON file contents. source node only links to its referer when every visit carrying that tag
came from the same origin. External exits are small nodes fanned outwards
from their source page), per-page view
counts, the top transitions and the 50 most recent visit trails. Data is
streamed live over `WebSocket /_api/ws/analytics`, which pushes the latest
JSON snapshot on connect and again whenever the analytics file is updated
(with a small server-side debounce to avoid flooding under high traffic).
+208 -111
View File
@@ -1,20 +1,23 @@
<script setup> <script setup>
// Analytics viewer rendered as a normal page inside #main. Fetches the raw // Analytics viewer rendered as a normal page inside #main. Receives live
// collected data from /_api/analytics (admin-gated by the auth proxy) and // analytics data over /_api/ws/analytics (admin-gated by the auth proxy) and
// renders totals, smoothed visit/views curves, a transition map, and recent // renders totals, smoothed visit/views curves, a transition map, and recent
// visit/crawler tables. Read-only. // visit/crawler tables. Read-only.
// See docs/analytics.md for the data format. // See docs/analytics.md for the data format.
import { computed, onMounted, ref, watch } from 'vue' import { computed, onMounted, onUnmounted, ref, watch } from 'vue'
import { RANGES } from './analytics/time.js' import { RANGES } from './analytics/time.js'
import { import {
calcReadStats,
calcTotalViews, calcTotalViews,
copyIp, copyIp,
countCrawlerUas, copyList,
formatCounts, formatCount,
formatAbuseRows,
formatCrawlerRows, formatCrawlerRows,
formatVisitRows, formatVisitRows,
} from './analytics/format.js' } from './analytics/format.js'
import * as flagSvgs from 'country-flag-icons/string/3x2' import TrailLink from './TrailLink.vue'
import VisitorCell from './VisitorCell.vue'
import TransitionGraph from './TransitionGraph.vue' import TransitionGraph from './TransitionGraph.vue'
import VisitorCharts from './VisitorCharts.vue' import VisitorCharts from './VisitorCharts.vue'
@@ -22,55 +25,78 @@ const props = defineProps({
initialRange: { type: String, default: 'week' }, initialRange: { type: String, default: 'week' },
}) })
const ABUSE_MAX_LINES = 5
const data = ref(null) const data = ref(null)
const pageTree = ref(null) const pageTree = ref(null)
const error = ref('') const error = ref('')
const now = ref(Date.now())
let ws = null
let reconnectTimeout = null
let timeInterval = null
onMounted(async () => { function connectAnalytics() {
try { if (ws) return
const res = await fetch('/_api/analytics') const proto = location.protocol === 'https:' ? 'wss:' : 'ws:'
if (!res.ok) throw new Error(res.statusText) ws = new WebSocket(`${proto}//${location.host}/_api/ws/analytics`)
data.value = await res.json() ws.onopen = () => { error.value = '' }
} catch { ws.onmessage = (event) => {
try {
data.value = JSON.parse(event.data)
} catch {
error.value = 'analytics data could not be loaded'
}
}
ws.onerror = () => {
error.value = 'analytics data could not be loaded' error.value = 'analytics data could not be loaded'
} }
ws.onclose = () => {
ws = null
reconnectTimeout = setTimeout(connectAnalytics, 2000)
}
}
onMounted(async () => {
connectAnalytics()
now.value = Date.now()
timeInterval = setInterval(() => { now.value = Date.now() }, 1000)
// The site tree for the transition map (all pages in menu order). Not // The site tree for the transition map (all pages in menu order). Not
// fatal: without it the map falls back to transition endpoints only. // fatal: without it the map just narrows to pages seen in transitions.
try { try {
const res = await fetch('/_api/pages') const res = await fetch('/_api/pages')
if (res.ok) pageTree.value = await res.json() if (res.ok) pageTree.value = await res.json()
} catch { /* map just narrows to pages seen in transitions */ } } catch { /* map just narrows to pages seen in transitions */ }
}) })
onUnmounted(() => {
if (reconnectTimeout) clearTimeout(reconnectTimeout)
if (timeInterval) clearInterval(timeInterval)
if (ws) {
ws.onclose = null
ws.close()
ws = null
}
})
const visits = computed(() => data.value?.visits || []) const visits = computed(() => data.value?.visits || [])
const totalViews = computed(() => calcTotalViews(data.value?.views)) const totalViews = computed(() => calcTotalViews(data.value?.views))
const readStats = computed(() => calcReadStats(visits.value))
const range = ref(RANGES[props.initialRange] ? props.initialRange : 'week') const range = ref(RANGES[props.initialRange] ? props.initialRange : 'week')
// Keep the URL shareable when the range changes. // Keep the URL shareable when the range changes.
watch(range, (r) => { watch(range, (r) => {
const url = new URL(location.href) const url = new URL(location.href)
url.searchParams.set('range', r) url.hash = r
history.replaceState(null, '', url) history.replaceState(null, '', url)
}) })
const visitRows = computed(() => formatVisitRows(visits.value, pageTree.value)) const clients = computed(() => data.value?.clients || {})
const visitRows = computed(() => formatVisitRows(visits.value, clients.value, pageTree.value, now.value))
const crawlers = computed(() => data.value?.crawlers || []) const crawlers = computed(() => data.value?.crawlers || [])
const crawlerRows = computed(() => formatCrawlerRows(crawlers.value)) const crawlerRows = computed(() => formatCrawlerRows(crawlers.value, clients.value, pageTree.value, now.value))
const topCrawlerUas = computed(() => countCrawlerUas(crawlers.value).slice(0, 10)) const abuseRows = computed(() => formatAbuseRows(data.value?.abuse || [], clients.value, now.value))
function flagSvg(code) {
return flagSvgs[code?.toUpperCase()] || ''
}
function countryName(code) {
if (!code) return ''
try {
return new Intl.DisplayNames(['en'], { type: 'region' }).of(code.toUpperCase())
} catch {
return ''
}
}
</script> </script>
<template> <template>
@@ -90,8 +116,10 @@ function countryName(code) {
<p v-else-if="!data" class="loading">loading</p> <p v-else-if="!data" class="loading">loading</p>
<template v-else> <template v-else>
<section class="totals"> <section class="totals">
<div><strong>{{ visits.length }}</strong> visits</div> <div><strong :title="String(visits.length)">{{ formatCount(visits.length) }}</strong> visits</div>
<div><strong>{{ totalViews }}</strong> page views</div> <div><strong :title="String(totalViews)">{{ formatCount(totalViews) }}</strong> page views</div>
<div><strong>{{ readStats.avgMinPerVisit }}</strong> min/visit</div>
<div><strong>{{ readStats.avgArticleMedianMin }}</strong> min article read</div>
</section> </section>
<VisitorCharts :data="data" :range="range" /> <VisitorCharts :data="data" :range="range" />
@@ -103,79 +131,112 @@ function countryName(code) {
<table class="visit-table"> <table class="visit-table">
<thead> <thead>
<tr> <tr>
<th>when</th>
<th>trail</th> <th>trail</th>
<th>referer</th> <th>visitor</th>
<th>ip</th> <th class="last-seen">last seen</th>
<th>lang</th>
<th>country</th>
<th>ua</th>
<th>utm</th>
</tr> </tr>
</thead> </thead>
<tbody> <tbody>
<tr v-for="(v, i) in visitRows" :key="i"> <tr v-for="(v, i) in visitRows" :key="i">
<td class="when">{{ v.when }}</td>
<td class="trail"> <td class="trail">
<a v-for="(s, si) in v.trail" :key="si" <TrailLink v-if="v.refererStep" :step="v.refererStep" @close="$emit('close')" />
:href="s.path" :title="s.title" @click="emit('close')"> <span v-if="v.utm && v.utm !== '—'" class="utm-tag small muted" :title="v.utmTitle">{{ v.utm }}</span>
{{ s.slug }} <TrailLink v-for="(s, si) in v.trail" :key="si" :step="s" @close="$emit('close')" />
</a>
</td> </td>
<td>{{ v.referer }}</td> <VisitorCell
<td> :ip="v.ip"
<span class="clickable-ip" :ip-display="v.ipDisplay"
:title="`Click to copy full IP: ${v.ip}`" :ua="v.ua"
@click="copyIp(v.ip)">{{ v.ipDisplay }}</span> :ua-raw="v.uaRaw"
</td> :country="v.country"
<td>{{ v.lang }}</td> :city="v.city"
<td class="country"> :lang="v.lang"
<span v-if="flagSvg(v.country)" class="flag" v-html="flagSvg(v.country)" :title="countryName(v.country) || v.country"></span> :lang-display="v.langDisplay"
<template v-else></template> :is-host="v.isHost"
</td> />
<td class="ua" :title="v.uaRaw">{{ v.ua }}</td> <td class="last-seen muted"
<td>{{ v.utm }}</td> :title="v.lastSeenLocal"
@click="copyList(v.lastSeenIso, $event)">{{ v.lastSeen }}</td>
</tr> </tr>
</tbody> </tbody>
</table> </table>
</div> </div>
<p v-else class="empty">no visits recorded yet</p> <p v-else class="empty">no visits recorded yet</p>
</section>
<section>
<h2>Crawlers</h2>
<div v-if="topCrawlerUas.length" class="crawler-top-uas">
<p><strong>top UAs:</strong> {{ formatCounts(topCrawlerUas) }}</p>
</div>
<div v-if="crawlerRows.length" class="visit-table-wrap"> <div v-if="crawlerRows.length" class="visit-table-wrap">
<table class="visit-table"> <table class="visit-table">
<thead> <thead>
<tr> <tr>
<th>when</th> <th>pages crawled</th>
<th>entry</th> <th>visitor</th>
<th>ip</th> <th class="last-seen">last seen</th>
<th>ua</th>
<th>referer</th>
<th>query</th>
</tr> </tr>
</thead> </thead>
<tbody> <tbody>
<tr v-for="(c, i) in crawlerRows" :key="i"> <tr v-for="(c, i) in crawlerRows" :key="i">
<td class="when">{{ c.when }}</td> <td class="trail">
<td>{{ c.entry }}</td> <TrailLink v-for="(s, si) in c.pages" :key="si" :step="s" :count="s.count" @close="$emit('close')" />
<td>
<span class="clickable-ip"
:title="`Click to copy full IP: ${c.ip}`"
@click="copyIp(c.ip)">{{ c.ipDisplay }}</span>
</td> </td>
<td class="ua" :title="c.uaRaw">{{ c.ua }}</td> <VisitorCell
<td>{{ c.referer }}</td> :ip="c.ip"
<td>{{ c.query }}</td> :ip-display="c.ipDisplay"
:ua="c.ua"
:ua-raw="c.uaRaw"
:country="c.country"
:city="c.city"
:lang="c.lang"
:lang-display="c.langDisplay"
:is-host="c.isHost"
/>
<td class="last-seen muted"
:title="c.lastSeenLocal"
@click="copyList(c.lastSeenIso, $event)">{{ c.lastSeen }}</td>
</tr> </tr>
</tbody> </tbody>
</table> </table>
</div> </div>
<p v-else class="empty">no crawler hits recorded yet</p> <p v-else class="empty">no crawler hits recorded yet</p>
<div v-if="abuseRows.length" class="visit-table-wrap">
<table class="visit-table">
<thead>
<tr>
<th>paths abused</th>
<th>visitor</th>
<th class="last-seen">last seen</th>
</tr>
</thead>
<tbody>
<tr v-for="(a, i) in abuseRows" :key="i">
<td class="trail abuse-list clickable-list"
@click="copyList(a.allPaths, $event)">
<div class="abuse-items">
<span v-for="(p, pi) in a.paths.slice(0, ABUSE_MAX_LINES)" :key="pi"
class="inline-item">
<small v-if="p.count > 1" class="muted">{{ formatCount(p.count) }}×</small>{{ p.path }}
</span>
<small v-if="a.paths.length > ABUSE_MAX_LINES" class="muted">+{{ a.paths.length - ABUSE_MAX_LINES }} more</small>
</div>
</td>
<VisitorCell
:ip="a.ip"
:ip-display="a.ipDisplay"
:ua="a.ua"
:ua-raw="a.uaRaw"
:country="a.country"
:city="a.city"
:lang="a.lang"
:lang-display="a.langDisplay"
:is-host="a.isHost"
:variant-count="a.clientCount"
/>
<td class="last-seen muted"
:title="a.lastSeenLocal"
@click="copyList(a.lastSeenIso, $event)">{{ a.lastSeen }}</td>
</tr>
</tbody>
</table>
</div>
</section> </section>
</template> </template>
</div> </div>
@@ -250,6 +311,15 @@ function countryName(code) {
margin-top: 1.8rem; margin-top: 1.8rem;
} }
.analytics-view a {
color: var(--text);
text-decoration: none;
}
.analytics-view a:hover { color: var(--accent); }
.analytics-view :deep(.muted) { color: var(--muted); }
.analytics-view :deep(.small) { font-size: 0.75em; }
.totals { .totals {
display: flex; display: flex;
gap: 2rem; gap: 2rem;
@@ -264,7 +334,6 @@ function countryName(code) {
.visit-table { .visit-table {
width: 100%; width: 100%;
border-collapse: collapse; border-collapse: collapse;
font-family: monospace;
font-size: 0.82rem; font-size: 0.82rem;
line-height: 1.3; line-height: 1.3;
} }
@@ -286,9 +355,11 @@ function countryName(code) {
background: var(--bg, Canvas); background: var(--bg, Canvas);
} }
.visit-table .when { .visit-table .last-seen {
width: 5rem;
text-align: right;
white-space: nowrap; white-space: nowrap;
color: var(--muted); cursor: pointer;
} }
.visit-table .trail { .visit-table .trail {
@@ -296,48 +367,74 @@ function countryName(code) {
overflow-wrap: break-word; overflow-wrap: break-word;
} }
.visit-table .trail a { .visit-table .trail a,
color: var(--text); .visit-table .trail-link {
text-decoration: none; display: inline-block;
max-width: 8rem;
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: bottom;
} }
.visit-table .trail a:hover { color: var(--accent); } .visit-table .trail > * + * {
.visit-table .trail a + a {
margin-left: 0.5rem; margin-left: 0.5rem;
} }
.visit-table .clickable-ip { .visit-table .utm-tag {
cursor: pointer; display: inline-block;
text-decoration: underline; max-width: 100%;
text-decoration-style: dotted; padding: 0.05rem 0.4rem;
} border: 1px solid var(--line);
border-radius: 0.25rem;
.visit-table .clickable-ip:hover { white-space: nowrap;
color: var(--accent);
}
.visit-table .ua {
max-width: 18rem;
overflow: hidden; overflow: hidden;
text-overflow: ellipsis; text-overflow: ellipsis;
vertical-align: bottom;
}
.visit-table .clickable-list {
cursor: pointer;
max-width: 22rem;
}
.visit-table .abuse-items {
display: flex;
flex-wrap: wrap;
gap: 0.15rem 0.5rem;
align-items: baseline;
}
.visit-table .inline-item {
max-width: 18rem;
min-width: 0;
white-space: nowrap; white-space: nowrap;
}
.visit-table .country .flag {
display: inline-flex;
width: 18px;
height: 12px;
border-radius: 2px;
overflow: hidden; overflow: hidden;
border: 1px solid var(--line); text-overflow: ellipsis;
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset; word-break: keep-all;
hyphens: none;
} }
.visit-table .country .flag :deep(svg) { .visit-table :deep(.clickable-ip),
width: 100%; .visit-table .clickable-list,
height: 100%; .visit-table .last-seen {
display: block; cursor: pointer;
position: relative;
}
.visit-table :deep(.copy-popup) {
position: absolute;
bottom: calc(100% + 0.25rem);
left: 50%;
transform: translateX(-50%);
padding: 0.15rem 0.4rem;
background: var(--text, CanvasText);
color: var(--bg, Canvas);
border-radius: 0.25rem;
font-size: 0.75rem;
white-space: nowrap;
pointer-events: none;
z-index: 10;
} }
.crawler-top-uas { .crawler-top-uas {
+22
View File
@@ -0,0 +1,22 @@
<script setup>
import { formatCount } from './analytics/format.js'
defineProps({
step: { type: Object, required: true },
count: { type: Number, default: 0 },
})
defineEmits(['close'])
</script>
<template>
<a class="trail-link"
:href="step.path"
:title="count > 1 ? `${step.title} (${count} hits)` : step.title"
:target="step.external ? '_blank' : undefined"
:rel="step.external ? 'noopener' : undefined"
@click="(e) => { if (!step.external) $emit('close') }">
<small v-if="count > 1" class="muted">{{ formatCount(count) }}×</small>
<span>{{ step.slug }}</span>
</a>
</template>
+79 -19
View File
@@ -7,7 +7,8 @@
* range, exactly like the charts and per-page views do. * range, exactly like the charts and per-page views do.
*/ */
import { computed, onBeforeUnmount, shallowRef, watch } from 'vue' import { computed, onBeforeUnmount, shallowRef, watch } from 'vue'
import { rangeWindow } from './analytics/time.js' import { rangeWindow, WEEK } from './analytics/time.js'
import { formatCount } from './analytics/format.js'
import { import {
TNODE_R, TNODE_R,
BEAD_R, BEAD_R,
@@ -15,6 +16,7 @@ import {
buildTransitionGraph, buildTransitionGraph,
filterTransitionsByRange, filterTransitionsByRange,
filterViewsByRange, filterViewsByRange,
filterVisitsByRange,
} from './analytics/transitions.js' } from './analytics/transitions.js'
const props = defineProps({ const props = defineProps({
@@ -25,6 +27,19 @@ const props = defineProps({
const window = computed(() => rangeWindow(props.range)) const window = computed(() => rangeWindow(props.range))
const visualScale = computed(() => {
const { t0, t1 } = window.value
if (t0 != null && t1 != null) return WEEK / (t1 - t0)
// 'all': scale by the actual data span.
const times = new Set()
for (const buckets of Object.values(props.data?.views || {})) {
for (const k of Object.keys(buckets)) times.add(Date.parse(k))
}
const arr = [...times]
if (arr.length < 2) return 1
return WEEK / (Math.max(...arr) - Math.min(...arr))
})
const filteredData = computed(() => { const filteredData = computed(() => {
if (!props.data) return null if (!props.data) return null
const { t0, t1 } = window.value const { t0, t1 } = window.value
@@ -34,9 +49,13 @@ const filteredData = computed(() => {
} }
}) })
const filteredVisits = computed(() =>
filterVisitsByRange(props.data?.visits, window.value.t0, window.value.t1),
)
const graph = computed(() => const graph = computed(() =>
filteredData.value filteredData.value
? buildTransitionGraph(filteredData.value, props.pageTree) ? buildTransitionGraph(filteredData.value, props.pageTree, filteredVisits.value, visualScale.value)
: null, : null,
) )
@@ -47,16 +66,24 @@ const graph = computed(() =>
const beads = shallowRef([]) const beads = shallowRef([])
let rafId = 0 let rafId = 0
const MAX_BEAD_RATE = 120 // upper bound on total beads per second
const startBeads = (flows) => { const startBeads = (flows) => {
cancelAnimationFrame(rafId) cancelAnimationFrame(rafId)
beads.value = [] beads.value = []
if (!flows?.length) return if (!flows?.length) return
if (matchMedia('(prefers-reduced-motion: reduce)').matches) return if (matchMedia('(prefers-reduced-motion: reduce)').matches) return
// Cap the total bead emission rate so a busy range cannot spawn enough
// beads to kill the page. Existing per-range time scaling is preserved;
// this is only a proportional emergency throttle when the limit is hit.
const totalRate = flows.reduce((s, f) => s + 1 / f.interval, 0)
const scale = totalRate > MAX_BEAD_RATE ? MAX_BEAD_RATE / totalRate : 1
const live = [] // { flow, t0 } — one entry per bead in flight const live = [] // { flow, t0 } — one entry per bead in flight
const now = performance.now() const now = performance.now()
const emitters = flows.map((flow) => { const emitters = flows.map((flow) => {
const interval = flow.interval * 1000 const interval = (flow.interval / scale) * 1000
// Pre-fill the traversal with evenly spaced beads (random phase), so // Pre-fill the traversal with evenly spaced beads (random phase), so
// the flow appears already running instead of starting empty. // the flow appears already running instead of starting empty.
const phase = Math.random() * interval const phase = Math.random() * interval
@@ -100,26 +127,56 @@ onBeforeUnmount(() => cancelAnimationFrame(rafId))
<section v-if="graph"> <section v-if="graph">
<svg class="tmap" :viewBox="`${graph.bounds.x0} ${graph.bounds.y0} ${graph.bounds.x1 - graph.bounds.x0} ${graph.bounds.y1 - graph.bounds.y0}`" <svg class="tmap" :viewBox="`${graph.bounds.x0} ${graph.bounds.y0} ${graph.bounds.x1 - graph.bounds.x0} ${graph.bounds.y1 - graph.bounds.y0}`"
role="img" aria-label="map of transitions between pages"> role="img" aria-label="map of transitions between pages">
<defs>
<!-- Unit-radius circle; only the portion near the bottom is used. -->
<path id="tnode-label-arc" d="M 0,-1 A 1,1 0 1,0 0,1 A 1,1 0 1,0 -0.001,-1" />
</defs>
<path v-for="(a, i) in graph.arcs" :key="'a' + i" <path v-for="(a, i) in graph.arcs" :key="'a' + i"
:d="a.d" class="tarc" /> :d="a.d" class="tarc" />
<path v-for="(e, i) in graph.edges" :key="'e' + i" <path v-for="(e, i) in graph.edges" :key="'e' + i"
:d="e.d" class="tconn"> :d="e.d" :class="['tconn', e.external && 'tconn-exit']">
<title>{{ e.title }}</title> <title>{{ e.title }}</title>
</path> </path>
<circle v-for="(b, i) in beads" :key="'b' + i" <circle v-for="(b, i) in beads" :key="'b' + i"
:cx="b.x" :cy="b.y" :r="BEAD_R" class="tbead" /> :cx="b.x" :cy="b.y" :r="BEAD_R" class="tbead" />
<g v-for="(x, i) in graph.extNodes" :key="'x' + i"> <g v-for="(x, i) in graph.extNodes" :key="'x' + i">
<circle :cx="x.x" :cy="x.y" :r="x.r" class="txnode"> <a v-if="x.href" :href="x.href" target="_blank" rel="noopener" :title="x.path">
<title>{{ x.path }}</title> <circle :cx="x.x" :cy="x.y" :r="x.r"
</circle> :class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
<text :x="x.x" :y="x.y + x.r + 11" class="txlabel">{{ x.label }}</text> <text :transform="`translate(${x.x}, ${x.y}) scale(${x.r - 4})`" class="tnodeslug" :style="{ '--node-r': x.r - 4 }">
<textPath href="#tnode-label-arc" startOffset="50%" text-anchor="middle" side="right">{{ x.label }}</textPath>
</text>
<text :x="x.x" :y="x.y + 4" class="tnodecount">{{ formatCount(x.count) }}</text>
</a>
<g v-else :title="x.path">
<circle :cx="x.x" :cy="x.y" :r="x.r"
:class="['txnode', x.kind === 'source' ? 'txnode-source' : 'txnode-exit']" />
<text :transform="`translate(${x.x}, ${x.y}) scale(${x.r - 4})`" class="tnodeslug" :style="{ '--node-r': x.r - 4 }">
<textPath href="#tnode-label-arc" startOffset="50%" text-anchor="middle" side="right">{{ x.label }}</textPath>
</text>
<text :x="x.x" :y="x.y + 4" class="tnodecount">{{ formatCount(x.count) }}</text>
</g>
</g> </g>
<g v-for="n in graph.nodes" :key="n.path"> <g v-for="n in graph.nodes" :key="n.path">
<a :href="n.path" :title="n.title"> <a v-if="!n.hidden" :href="n.path" :title="n.title">
<circle :cx="n.x" :cy="n.y" :r="TNODE_R" class="tnode" /> <circle :cx="n.x" :cy="n.y" :r="TNODE_R" class="tnode" />
<text :x="n.x" :y="n.y - 2" class="tnodeslug">{{ n.label }}</text> <text :transform="`translate(${n.x}, ${n.y}) scale(${TNODE_R - 4})`" class="tnodeslug" :style="{ '--node-r': TNODE_R - 4 }">
<text :x="n.x" :y="n.y + 12" class="tnodecount">{{ n.views }}</text> <textPath href="#tnode-label-arc" startOffset="50%" text-anchor="middle" side="right">{{ n.label }}</textPath>
</text>
<text :x="n.x" :y="n.y + 4" class="tnodecount">
{{ n.readMin ? `${formatCount(n.views)}×${n.readMin}m` : formatCount(n.views) }}
</text>
</a> </a>
<template v-else>
<text :x="n.x" :y="n.y"
:transform="`rotate(${n.angle * 180 / Math.PI}, ${n.x}, ${n.y})`"
class="tnodehidden" text-anchor="start" dominant-baseline="middle"></text>
<text :x="n.x + Math.cos(n.angle) * 10"
:y="n.y + Math.sin(n.angle) * 10"
:transform="`rotate(${(n.angle + (Math.cos(n.angle) < 0 ? Math.PI : 0)) * 180 / Math.PI}, ${n.x + Math.cos(n.angle) * 10}, ${n.y + Math.sin(n.angle) * 10})`"
:text-anchor="Math.cos(n.angle) < 0 ? 'end' : 'start'"
class="tnodehidden" dominant-baseline="middle">{{ n.label }}</text>
</template>
</g> </g>
</svg> </svg>
</section> </section>
@@ -137,6 +194,9 @@ onBeforeUnmount(() => cancelAnimationFrame(rafId))
fill: var(--accent); fill: var(--accent);
opacity: 0.4; /* uniform, not strength-encoded: width carries that */ opacity: 0.4; /* uniform, not strength-encoded: width carries that */
} }
.tmap .tconn-exit {
fill: var(--text);
}
.tmap .tbead { .tmap .tbead {
fill: var(--accent); fill: var(--accent);
opacity: 0.85; opacity: 0.85;
@@ -144,14 +204,10 @@ onBeforeUnmount(() => cancelAnimationFrame(rafId))
} }
.tmap .txnode { .tmap .txnode {
fill: var(--bg, Canvas); fill: var(--bg, Canvas);
stroke: var(--muted); stroke-width: 1.5;
stroke-width: 1;
}
.tmap .txlabel {
fill: var(--muted);
font-size: 9px;
text-anchor: middle;
} }
.tmap .txnode-source { stroke: var(--text); }
.tmap .txnode-exit { stroke: var(--text); }
.tmap .tarc { .tmap .tarc {
fill: none; fill: none;
stroke: var(--line); stroke: var(--line);
@@ -164,7 +220,7 @@ onBeforeUnmount(() => cancelAnimationFrame(rafId))
} }
.tmap .tnodeslug { .tmap .tnodeslug {
fill: var(--text); fill: var(--text);
font-size: 11px; font-size: calc(11px / var(--node-r, 34));
text-anchor: middle; text-anchor: middle;
} }
.tmap a { cursor: pointer; } .tmap a { cursor: pointer; }
@@ -174,6 +230,10 @@ onBeforeUnmount(() => cancelAnimationFrame(rafId))
font-size: 10px; font-size: 10px;
text-anchor: middle; text-anchor: middle;
} }
.tmap .tnodehidden {
fill: var(--text);
font-size: 9px;
}
section { margin-top: 1.8rem; } section { margin-top: 1.8rem; }
</style> </style>
+156
View File
@@ -0,0 +1,156 @@
<script setup>
// Visitor metadata cell shared by the recent-visits, crawlers, and abuse tables.
// Displays IP/network/host, country flag/city, UA, and language when available.
// Clicking the IP copies the full address to the clipboard.
// ``variantCount`` overrides the UA line to warn when multiple client
// fingerprints share the same IP (e.g. a scanner rotating UAs).
import { computed } from 'vue'
import * as flagSvgs from 'country-flag-icons/string/3x2'
import { copyIp, formatLang } from './analytics/format.js'
const props = defineProps({
ip: { type: String, default: '' },
ipDisplay: { type: String, default: '—' },
ua: { type: String, default: '' },
uaRaw: { type: String, default: '' },
country: { type: String, default: '' },
city: { type: String, default: '' },
lang: { type: String, default: '' },
langDisplay: { type: String, default: '' },
isHost: { type: Boolean, default: false },
variantCount: { type: Number, default: 1 },
})
const hasCountry = computed(() => !!(props.country && props.country !== '—'))
const hasCity = computed(() => !!(props.city && props.city !== '—'))
const hasLocale = computed(() => hasCountry.value || hasCity.value)
const langValue = computed(() => props.langDisplay || formatLang(props.lang))
const showLang = computed(() => langValue.value && langValue.value !== '—')
function flagSvg(code) {
return flagSvgs[code?.toUpperCase()] || ''
}
function countryName(code) {
if (!code) return ''
try {
return new Intl.DisplayNames(['en'], { type: 'region' }).of(code.toUpperCase())
} catch {
return ''
}
}
</script>
<template>
<td class="visitor-cell" :class="{ 'host-cell': isHost }">
<div class="visitor-rows">
<div class="visitor-row">
<div class="locale-line">
<span v-if="flagSvg(country)" class="flag" v-html="flagSvg(country)" :title="countryName(country) || country"></span>
<template v-if="hasCity"><small class="city-name muted">{{ city }}</small></template>
<template v-else-if="!hasLocale"></template>
</div>
<div class="ip-line">
<span class="clickable-ip small muted"
:title="ip"
@click="copyIp(ip, $event)">{{ ipDisplay }}</span>
</div>
</div>
<div class="visitor-row">
<div class="ua-line">
<small v-if="variantCount > 1" class="muted variant-hint">{{ variantCount }} client variations</small>
<small v-else class="muted" :title="uaRaw">{{ ua || '—' }}</small>
</div>
<div v-if="showLang && variantCount <= 1" class="locale-lang"><small class="muted">{{ langValue }}</small></div>
</div>
</div>
</td>
</template>
<style scoped>
.visitor-cell {
width: 18em;
max-width: 18em;
overflow: hidden;
text-overflow: ellipsis;
vertical-align: top;
}
.visitor-cell.host-cell {
text-align: right;
}
.visitor-rows {
display: flex;
flex-direction: column;
gap: 0.15rem;
}
.visitor-row {
display: flex;
align-items: center;
justify-content: space-between;
gap: 0.5rem;
}
.visitor-row > * {
min-width: 0;
}
.locale-line,
.ip-line,
.ua-line {
flex: 1 1 auto;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
}
.locale-line {
text-align: left;
display: flex;
align-items: center;
gap: 0.3rem;
}
.ip-line {
text-align: right;
}
.ua-line {
text-align: left;
}
.locale-lang {
flex: 0 0 auto;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
text-align: right;
}
.city-name {
display: inline-block;
overflow: hidden;
text-overflow: ellipsis;
white-space: nowrap;
vertical-align: middle;
}
.flag {
display: inline-flex;
width: 18px;
height: 12px;
border-radius: 2px;
overflow: hidden;
border: 1px solid var(--line);
box-shadow: 0 0 0 1px rgba(0, 0, 0, 0.2) inset;
vertical-align: middle;
}
.flag :deep(svg) {
width: 100%;
height: 100%;
display: block;
}
</style>
+42 -15
View File
@@ -2,10 +2,12 @@
/** /**
* Visitor and page-view smoothed curves for a single shared time range. * Visitor and page-view smoothed curves for a single shared time range.
*/ */
import { computed } from 'vue' import { computed, onMounted, onUnmounted, ref } from 'vue'
import { makeSeries } from './analytics/time.js' import { makeSeries } from './analytics/time.js'
import { CHART_H, CHART_W, buildChart, fmtY } from './analytics/chart.js' import { CHART_H, CHART_W, buildChart, fmtY } from './analytics/chart.js'
const DAY_REFRESH_MS = 15000
const props = defineProps({ const props = defineProps({
data: { type: Object, default: null }, data: { type: Object, default: null },
range: { type: String, required: true }, range: { type: String, required: true },
@@ -22,33 +24,52 @@ const allViews = computed(() => {
const visitSeries = computed(() => makeSeries(props.data?.site_visits, props.range)) const visitSeries = computed(() => makeSeries(props.data?.site_visits, props.range))
const viewSeries = computed(() => makeSeries(allViews.value, props.range)) const viewSeries = computed(() => makeSeries(allViews.value, props.range))
const unit = computed(() => (props.range === 'week' ? 'h' : 'day'))
const visitChart = computed(() => buildChart(visitSeries.value)) function freqLabel(unit) {
const viewChart = computed(() => buildChart(viewSeries.value)) return unit === '5min' ? '5 min' : unit === 'hour' ? 'hourly' : 'daily'
}
const now = ref(Date.now())
let refreshInterval = null
onMounted(() => {
refreshInterval = setInterval(() => { now.value = Date.now() }, DAY_REFRESH_MS)
})
onUnmounted(() => {
if (refreshInterval) clearInterval(refreshInterval)
})
const visitChart = computed(() => buildChart(visitSeries.value, now.value))
const viewChart = computed(() => buildChart(viewSeries.value, now.value))
</script> </script>
<template> <template>
<section v-for="c in [ <section v-for="c in [
{ ylabel: 'visitors', chart: visitChart, empty: 'no visits recorded yet' }, { ylabel: 'visits', chart: visitChart, empty: 'no visits recorded yet' },
{ ylabel: 'views', chart: viewChart, empty: 'no views recorded yet' }, { ylabel: 'views', chart: viewChart, empty: 'no views recorded yet' },
]" :key="c.ylabel"> ]" :key="c.ylabel">
<template v-if="c.chart"> <template v-if="c.chart">
<div class="chartwrap"> <div class="chartwrap">
<div class="plot"> <div class="plot">
<div class="plotarea"> <div class="plotarea">
<span class="yaxis-label">{{ c.ylabel }}/{{ unit }}</span> <span class="yaxis-label">{{ freqLabel(c.chart.unit) }} {{ c.ylabel }}</span>
<svg class="chart" :viewBox="`0 0 ${CHART_W} ${CHART_H}`" <svg class="chart" :viewBox="`0 0 ${CHART_W} ${CHART_H}`"
preserveAspectRatio="none" role="img" :aria-label="`${c.ylabel} per ${unit}`"> preserveAspectRatio="none" role="img" :aria-label="`${freqLabel(c.chart.unit)} ${c.ylabel}`">
<line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value" <line v-for="g in c.chart.majors.slice(1)" :key="'j' + g.value"
:x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" /> :x1="0" :x2="CHART_W" :y1="g.y" :y2="g.y" class="major" />
<template v-for="t in c.chart.xticks" :key="'t' + t.x"> <template v-for="t in c.chart.xticks" :key="'t' + t.x">
<line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H" <line v-if="t.line" :x1="t.x" :x2="t.x" :y1="0" :y2="CHART_H"
class="minor vertical" /> class="minor vertical" />
</template> </template>
<template v-for="(s, i) in c.chart.series" :key="i"> <template v-if="c.chart.bars">
<path v-if="s.area" :d="s.area" class="area" /> <rect v-for="(b, i) in c.chart.bars" :key="'b' + i"
<path :d="s.line" class="line" :style="{ opacity: s.opacity }" /> :x="b.x" :y="b.y" :width="b.width" :height="b.height" class="bar" />
<path :d="c.chart.skyline" class="line" />
</template>
<template v-else>
<template v-for="(s, i) in c.chart.series" :key="i">
<path v-if="s.area" :d="s.area" class="area" />
<path :d="s.line" class="line" :style="{ opacity: s.opacity }" />
</template>
</template> </template>
<line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5" <line :x1="0" :x2="CHART_W" :y1="CHART_H - 0.5" :y2="CHART_H - 0.5"
class="axis" /> class="axis" />
@@ -62,7 +83,7 @@ const viewChart = computed(() => buildChart(viewSeries.value))
</div> </div>
</div> </div>
</div> </div>
<div v-if="c.chart.series.length > 1" class="legend"> <div v-if="c.chart.series && c.chart.series.length > 1" class="legend">
<span v-for="(s, i) in c.chart.series" :key="i" :style="{ opacity: s.opacity }"> <span v-for="(s, i) in c.chart.series" :key="i" :style="{ opacity: s.opacity }">
● {{ s.label }} ● {{ s.label }}
</span> </span>
@@ -76,7 +97,7 @@ const viewChart = computed(() => buildChart(viewSeries.value))
/* The svg is stretched (preserveAspectRatio none), so all text lives in /* The svg is stretched (preserveAspectRatio none), so all text lives in
HTML overlays positioned by the same fractions the geometry uses. */ HTML overlays positioned by the same fractions the geometry uses. */
.chartwrap { .chartwrap {
padding-left: 2.2rem; /* y labels */ padding-left: 2.8rem; /* y labels */
} }
.plot { .plot {
@@ -103,8 +124,8 @@ const viewChart = computed(() => buildChart(viewSeries.value))
.ylab { .ylab {
position: absolute; position: absolute;
left: -2.2rem; left: -2.8rem;
width: 1.9rem; width: 2.6rem;
text-align: right; text-align: right;
transform: translateY(50%); transform: translateY(50%);
font-size: 0.7rem; font-size: 0.7rem;
@@ -154,6 +175,11 @@ const viewChart = computed(() => buildChart(viewSeries.value))
opacity: 0.15; opacity: 0.15;
} }
.chart .bar {
fill: var(--accent);
opacity: 0.15;
}
.chart .line { .chart .line {
fill: none; fill: none;
stroke: var(--accent); stroke: var(--accent);
@@ -176,10 +202,11 @@ const viewChart = computed(() => buildChart(viewSeries.value))
.yaxis-label { .yaxis-label {
position: absolute; position: absolute;
top: 50%; top: 50%;
left: -2.2rem; left: -2.8rem;
font-size: 0.7rem; font-size: 0.7rem;
color: var(--muted); color: var(--muted);
writing-mode: vertical-rl; writing-mode: vertical-rl;
white-space: nowrap;
transform: translateY(-50%) rotate(180deg); transform: translateY(-50%) rotate(180deg);
} }
+1 -1
View File
@@ -9,7 +9,7 @@ let app = null
export function mount(container) { export function mount(container) {
if (app) return if (app) return
app = createApp(AnalyticsView, { app = createApp(AnalyticsView, {
initialRange: new URLSearchParams(location.search).get('range') || 'week', initialRange: location.hash.slice(1) || 'week',
}) })
app.mount(container) app.mount(container)
} }
+98 -12
View File
@@ -5,7 +5,8 @@
* rates (hour on the week view, day on month+). * rates (hour on the week view, day on month+).
*/ */
import { DAY, HOUR, WEEK, mondayUTC } from './time.js' import { DAY, HOUR, MIN5, WEEK, mondayUTC } from './time.js'
import { formatCount } from './format.js'
export const CHART_W = 720 export const CHART_W = 720
export const CHART_H = 180 export const CHART_H = 180
@@ -14,9 +15,9 @@ export const PAD_TOP = 14 // room above the highest point
/** /**
* Y always starts at 0; the max is a multiple of a 1-2-5 major step with at * Y always starts at 0; the max is a multiple of a 1-2-5 major step with at
* most 5 intervals, so labeled ticks are always round and evenly divided. * most 5 intervals, so labeled ticks are always round and evenly divided.
* Values are per-unit rates, so small scales are legitimate (a lone visit * A minimum range of 10 keeps tiny near-zero values (e.g. a single visit)
* smoothes to well under 1/unit) — the floor is 1, not 10. Minor lines * from being enlarged to a fractional scale; minor lines subdivide each
* subdivide each major step in five when that yields integers. * major step in five when that yields integers.
*/ */
export function yScale(maxValue) { export function yScale(maxValue) {
let step = 1 let step = 1
@@ -27,9 +28,9 @@ export function yScale(maxValue) {
} }
} }
let max = Math.ceil(maxValue / step) * step let max = Math.ceil(maxValue / step) * step
if (max < 1) { if (max < 10) {
max = 1 max = 10
step = 0.5 step = 2
} }
const minor = step >= 5 && step % 5 === 0 ? step / 5 : null const minor = step >= 5 && step % 5 === 0 ? step / 5 : null
return { max, step, minor } return { max, step, minor }
@@ -206,9 +207,10 @@ export function spline(pts) {
} }
/** Build a full chart model from a series descriptor produced by time.js. */ /** Build a full chart model from a series descriptor produced by time.js. */
export function buildChart(input) { export function buildChart(input, now = Date.now()) {
if (!input || !input.series.length) return null if (!input || !input.series.length) return null
const { series, t0, t1, rate, binMinutes, unitMinutes } = input if (input.unit === '5min') return buildDayChart(input, now)
const { series, t0, t1, rate, binMinutes, unitMinutes, unit } = input
// Values are per-unit rates (hour on the week view, day on month+); the // Values are per-unit rates (hour on the week view, day on month+); the
// y max is derived from the *smoothed* curves so random single-bucket // y max is derived from the *smoothed* curves so random single-bucket
// spikes don't blow up the scale. Smoothing works on raw counts (its edge // spikes don't blow up the scale. Smoothing works on raw counts (its edge
@@ -283,7 +285,91 @@ export function buildChart(input) {
x: x(t), left: ((t - t0) / (t1 - t0)) * 100, x: x(t), left: ((t - t0) / (t1 - t0)) * 100,
label: fmtTick(t, t1 - t0), line: true, label: fmtTick(t, t1 - t0), line: true,
})) }))
return { max, majors, minors, series: drawn, xticks } return { max, majors, minors, series: drawn, xticks, unit }
}
/**
* Day view: 5-minute bars for the last 24 hours. Bars are drawn at raw
* counts; the skyline uses a projected full-bucket value for the still-open
* final bucket. The y scale is derived from the projected skyline maximum.
*/
export function buildDayChart(input, now = Date.now()) {
const { series, t0, t1 } = input
const points = series[0]?.points || []
const n = points.length
if (!n) return null
const bucketMs = (t1 - t0) / n
const bucketWidth = CHART_W / n
const gap = 0.2
const barWidth = Math.max(0.2, bucketWidth - gap)
const x = (i) => i * bucketWidth + gap / 2
const prevRaw = n > 1 ? points[n - 2].count : 0
const projected = points.map((p, i) => {
if (i !== n - 1) return p.count
const bucketStart = t0 + i * bucketMs
const elapsed = Math.max(1, Math.min(bucketMs, now - bucketStart))
// Blend the observed partial bucket with the previous full bucket:
// the longer the current bucket has run, the less we borrow from it.
const share = elapsed / bucketMs
return p.count + prevRaw * (1 - share)
})
const highest = Math.max(0, ...projected)
const { max, step, minor } = yScale(highest)
const y = (v) => PAD_TOP + (1 - Math.max(0, v) / max) * (CHART_H - PAD_TOP)
const bars = points.map((p, i) => {
const bx = x(i)
const by = y(p.count)
return {
x: bx,
y: by,
width: barWidth,
height: CHART_H - by,
raw: p.count,
projected: projected[i],
}
})
let skyline = ''
for (let i = 0; i < bars.length; i++) {
const b = bars[i]
const top = y(b.projected)
if (i === 0) {
skyline += `M${b.x},${top} H${b.x + b.width}`
} else {
skyline += ` V${top} H${b.x + b.width}`
}
}
const majors = []
const minors = []
const nMajor = Math.round(max / step)
for (let k = 0; k <= nMajor; k++) {
const v = k * step
majors.push({ value: v, y: y(v), bottom: (1 - PAD_TOP / CHART_H) * (v / max) * 100 })
}
if (minor) {
for (let v = minor; v < max; v += minor) {
if (v % step !== 0) minors.push({ y: y(v) })
}
}
const xticks = []
const tickStep = 3 * HOUR
const firstTick = Math.ceil(t0 / tickStep) * tickStep
for (let t = firstTick; t < t1; t += tickStep) {
if (t < t0) continue
const d = new Date(t)
xticks.push({
x: ((t - t0) / (t1 - t0)) * CHART_W,
left: ((t - t0) / (t1 - t0)) * 100,
label: `${String(d.getUTCHours()).padStart(2, '0')}:00`,
line: false,
})
}
return { bars, skyline: skyline.trim(), max, majors, minors, xticks, unit: '5min', series: [] }
} }
/** X ticks for year/all: Monday boundaries up to a quarter, UTC month /** X ticks for year/all: Monday boundaries up to a quarter, UTC month
@@ -327,7 +413,7 @@ export function fmtTick(t, span) {
return d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' }) return d.toLocaleDateString(undefined, { year: 'numeric', timeZone: 'UTC' })
} }
/** Y labels: integers when the step allows, one decimal for fractional steps. */ /** Y labels use the same compact formatter as text labels. */
export function fmtY(v) { export function fmtY(v) {
return Number.isInteger(v) ? String(v) : v.toFixed(1) return formatCount(v)
} }
+394 -48
View File
@@ -25,11 +25,44 @@ export const hostIP = (ip) => {
} }
} }
/** Copy the full IP to the clipboard, ignoring failures. */ function showCopiedFeedback(el) {
export async function copyIp(ip) { if (!el || typeof document === 'undefined') return
const popup = document.createElement('span')
popup.textContent = 'Copied!'
popup.className = 'copy-popup'
popup.style.cssText =
'position:absolute;bottom:calc(100% + 0.25rem);left:50%;' +
'transform:translateX(-50%);padding:0.15rem 0.4rem;' +
'background:var(--text, CanvasText);color:var(--bg, Canvas);' +
'border-radius:0.25rem;font-size:0.75rem;white-space:nowrap;' +
'pointer-events:none;z-index:10;'
el.classList.add('has-copy-popup')
el.appendChild(popup)
setTimeout(() => {
popup.remove()
el.classList.remove('has-copy-popup')
}, 1200)
}
/** Copy the full IP to the clipboard and show a brief "Copied!" popup. */
export async function copyIp(ip, event) {
if (!ip) return if (!ip) return
const el = event?.currentTarget
try { try {
await navigator.clipboard.writeText(ip) await navigator.clipboard.writeText(ip)
showCopiedFeedback(el)
} catch {
/* ignore */
}
}
/** Copy arbitrary text to the clipboard and show a brief "Copied!" popup. */
export async function copyList(text, event) {
if (!text) return
const el = event?.currentTarget
try {
await navigator.clipboard.writeText(text)
showCopiedFeedback(el)
} catch { } catch {
/* ignore */ /* ignore */
} }
@@ -44,6 +77,44 @@ export function calcTotalViews(views) {
return n return n
} }
// Very short reads are navigation/skims, not real reading time.
export const MIN_READ_SECONDS = 10
/** Average minutes per visit and average of per-article median read minutes. */
export function calcReadStats(visits) {
const perArticle = {}
let totalVisitSeconds = 0
let visitCount = 0
for (const v of visits || []) {
const secs = Object.values(v.read || {}).filter((s) => s >= MIN_READ_SECONDS)
if (!secs.length) continue
visitCount++
totalVisitSeconds += secs.reduce((a, b) => a + b, 0)
for (const [path, s] of Object.entries(v.read || {})) {
if (s >= MIN_READ_SECONDS) {
;(perArticle[path] || (perArticle[path] = [])).push(s)
}
}
}
const avgMinPerVisit = visitCount
? Math.max(1, Math.round(totalVisitSeconds / visitCount / 60))
: 0
let articleMedianSum = 0
const articleCount = Object.keys(perArticle).length
for (const arr of Object.values(perArticle)) {
arr.sort((a, b) => a - b)
const mid = Math.floor(arr.length / 2)
const median = arr.length % 2 ? arr[mid] : (arr[mid - 1] + arr[mid]) / 2
articleMedianSum += Math.max(MIN_READ_SECONDS, median)
}
const avgArticleMedianMin = articleCount
? Math.max(1, Math.round(articleMedianSum / articleCount / 60))
: 0
return { avgMinPerVisit, avgArticleMedianMin }
}
/** Build a path -> page title lookup from the site tree. */ /** Build a path -> page title lookup from the site tree. */
function buildTitleMap(pageTree) { function buildTitleMap(pageTree) {
const titles = new Map() const titles = new Map()
@@ -59,13 +130,130 @@ function buildTitleMap(pageTree) {
/** Last path segment for display; front page becomes a house icon. */ /** Last path segment for display; front page becomes a house icon. */
function slugOf(path) { function slugOf(path) {
return path === '/' ? '🏠' : path.split('/').pop() return path === '/' ? '🏠' : path.split('/').pop()
}
/** Host name of an external https origin, with scheme and www. stripped. */
function externalSlug(origin) {
try {
return new URL(origin).host.replace(/^www\./, '')
} catch {
return origin.replace(/^https?:\/\//, '').replace(/^www\./, '')
}
}
/** Format one trail step: an internal page or an external https origin. */
function stepOf(path, titles) {
if (path?.startsWith('/')) {
return { path, slug: slugOf(path), title: titles.get(path) || '', external: false, home: path === '/' }
}
if (path?.startsWith('https://')) {
return {
path,
slug: externalSlug(path),
title: 'External site',
external: true,
}
}
return null
}
/**
* Human-readable relative timestamp. Adapted from cista-storage: uses
* ``Intl.RelativeTimeFormat`` for short intervals and a compact date for
* anything older than a week.
*/
export function formatWhen(ts, now = Date.now()) {
const date = new Date(ts)
const diff = date.getTime() - now
const adiff = Math.abs(diff)
const formatter = new Intl.RelativeTimeFormat('en', { numeric: 'auto' })
if (adiff <= 5000) return 'now'
if (adiff <= 60000) {
return formatter
.format(Math.round(diff / 1000), 'second')
.replace(' ago', '')
.replaceAll(' ', '\u202F')
}
if (adiff <= 3600000) {
return formatter
.format(Math.round(diff / 60000), 'minute')
.replace('utes', '')
.replace('ute', '')
.replaceAll(' ', '\u202F')
}
if (adiff <= 86400000) {
return formatter
.format(Math.round(diff / 3600000), 'hour')
.replaceAll(' ', '\u202F')
}
if (adiff <= 604800000) {
return formatter
.format(Math.round(diff / 86400000), 'day')
.replaceAll(' ', '\u202F')
}
let d = date
.toLocaleDateString('en-ie', {
weekday: 'short',
year: 'numeric',
month: 'short',
day: 'numeric',
})
.replace('Sept', 'Sep')
if (d.length === 14) d = d.replace(' ', ' \u2007')
d = d.replaceAll(' ', '\u202F').replace('\u202F', '\u00A0')
d = d.slice(0, -4) + d.slice(-2)
return d
}
/** Full UTC timestamp for tooltips, e.g. "2026-08-21 00:20:48 UTC". */
export function formatWhenTooltip(ts) {
return new Date(ts).toISOString().replace('T', ' ').replace('Z', ' UTC')
}
/** Full local timestamp for tooltips, e.g. "21 Aug 2026, 17:38:48". */
export function formatWhenLocal(ts) {
return new Date(ts).toLocaleString('en-ie', {
year: 'numeric',
month: 'short',
day: 'numeric',
hour: '2-digit',
minute: '2-digit',
second: '2-digit',
})
}
/** Preserve locale case with the region/country subtag upper-cased. */
export function formatLang(value) {
if (!value || value === '—') return value
const parts = value.split('-')
if (parts.length > 1) {
parts[parts.length - 1] = parts[parts.length - 1].toUpperCase()
}
return parts.join('-')
}
/** ISO 8601 UTC timestamp without subseconds, e.g. "2026-08-21T00:20:48Z". */
export function formatWhenIso(ts) {
return `${new Date(ts).toISOString().split('.')[0]}Z`
}
/**
* Compact visitor counts: plain below 1k, then 1.2k / 10k / 1.2M.
* Truncated, not rounded.
*/
export function formatCount(n) {
if (n < 1000) return String(n)
if (n < 10000) return `${Math.trunc(n / 1000)}.${Math.trunc((n % 1000) / 100)}k`
if (n < 1_000_000) return `${Math.trunc(n / 1000)}k`
return `${Math.trunc(n / 1_000_000)}.${Math.trunc((n % 1_000_000) / 100_000)}M`
} }
/** /**
* Format recent visits for display, newest first. Each step is a linked slug * Format recent visits for display, newest first. Each step is a linked slug
* pointing to its article; external referers/origins and direct entries are * pointing to its article; external referers/origins are shown as their
* omitted. The link title shows the article heading when known. * domain name with the full origin as the link href. The link title shows the
* article heading when known, or "External site" for origins.
*/ */
export function formatRecentVisits(visits, pageTree, limit = 50) { export function formatRecentVisits(visits, pageTree, limit = 50) {
const titles = buildTitleMap(pageTree) const titles = buildTitleMap(pageTree)
@@ -73,13 +261,9 @@ export function formatRecentVisits(visits, pageTree, limit = 50) {
.reverse() .reverse()
.map((v) => ({ .map((v) => ({
when: new Date(v.start).toLocaleString(), when: new Date(v.start).toLocaleString(),
steps: [v.entry, ...(v.trail || [])] steps: [v.referer, v.entry, ...(v.trail || [])]
.filter((p) => p?.startsWith('/')) .map((p) => stepOf(p, titles))
.map((p) => ({ .filter(Boolean),
path: p,
slug: slugOf(p),
title: titles.get(p) || '',
})),
})) }))
.filter((v) => v.steps.length) .filter((v) => v.steps.length)
.slice(0, limit) .slice(0, limit)
@@ -121,66 +305,228 @@ export function formatCounts(entries) {
/** /**
* Count distinct User-Agent strings among crawler hits, most common first. * Count distinct User-Agent strings among crawler hits, most common first.
* Returns an array of [ua, count] pairs. * Returns an array of [ua, count] pairs. ``clients`` maps client hashes to
* client records.
*/ */
export function countCrawlerUas(crawlers) { export function countCrawlerUas(crawlers, clients) {
const counts = {} const counts = {}
for (const c of crawlers || []) { for (const c of crawlers || []) {
const value = c.ua_pretty || c.ua || '(no UA)' const client = (clients || {})[c.client] || {}
const value = client.ua_pretty || client.ua || '(no UA)'
counts[value] = (counts[value] || 0) + 1 counts[value] = (counts[value] || 0) + 1
} }
return Object.entries(counts).sort((a, b) => b[1] - a[1]) return Object.entries(counts).sort((a, b) => b[1] - a[1])
} }
/** /**
* Format raw crawler hit records as rows for a technical table. Missing * Reduce a reverse-DNS hostname to its right-most components that fit
* values become "—". * within ``limit`` characters. This keeps the meaningful main domain
* while avoiding absurdly long subdomains like ``xxx.yyy.zzz...provider.net``.
*/ */
export function formatCrawlerRows(crawlers) { export function mainDomain(host, limit = 24) {
const dash = (s) => (s || '—') if (!host) return host
return [...(crawlers || [])].reverse().map((c) => ({ const labels = host.split('.').filter(Boolean)
when: new Date(c.start).toLocaleString(), if (!labels.length) return host
entry: dash(c.entry), const parts = [labels.pop()]
ip: c.ip || '', while (labels.length) {
ipDisplay: c.host || hostIP(c.ip) || c.ip || '—', const next = labels[labels.length - 1]
ua: c.ua_pretty || c.ua || '', const candidate = `${next}.${parts.join('.')}`
uaRaw: c.ua || '', if (candidate.length > limit) break
referer: dash(c.referer), parts.unshift(labels.pop())
query: dash(c.query), }
})) return parts.join('.')
}
/**
* Group raw crawler hits by client hash and format each group as a row showing
* every internal page that crawler visited. Rows are sorted by total hits,
* most active crawler first, rather than by most recent hit.
* ``clients`` maps client hashes to client records.
*/
export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree)
const groups = new Map()
for (const c of crawlers || []) {
const client = (clients || {})[c.client] || {}
const g = groups.get(c.client) || {
clientHash: c.client,
client,
lastStart: 0,
pages: new Map(),
}
const start = new Date(c.start).getTime()
if (start > g.lastStart) g.lastStart = start
if (c.entry?.startsWith('/')) {
g.pages.set(c.entry, (g.pages.get(c.entry) || 0) + 1)
}
groups.set(c.client, g)
}
const totalHits = (g) => {
let n = 0
for (const c of g.pages.values()) n += c
return n
}
return [...groups.values()]
.sort((a, b) => totalHits(b) - totalHits(a) || b.lastStart - a.lastStart)
.slice(0, 10)
.map((g) => {
const client = g.client || {}
const host = client.host || ''
const isHost = !!host
return {
lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart),
lastSeenLocal: formatWhenLocal(g.lastStart),
pages: [...g.pages.entries()]
.sort((a, b) => b[1] - a[1])
.map(([path, count]) => ({ ...stepOf(path, titles), count })),
ip: client.ip || '',
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
isHost,
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
lang: client.lang || '—',
langDisplay: formatLang(client.lang),
country: client.country || '—',
city: client.city || '—',
total: totalHits(g),
}
})
}
/**
* Group abuse hits by IP and format each group as a row with the full paths
* probed. Identical paths are collapsed into one entry with their hit count.
* Flagged paths (the ones that triggered abuse classification) are lifted to
* the top, followed by other 404s, then document GETs from the abuser. Within
* each category paths are sorted by count descending, then earliest first.
* Rows are sorted by most recent hit first. Visitor metadata comes from the
* latest client hash seen for the IP; ``clientCount`` tells the visitor cell
* how many distinct client variations the IP produced. Paths are shown
* verbatim (query string included), not resolved against the page tree.
* ``clients`` maps client hashes to client records.
*/
export function formatAbuseRows(abuse, clients, now = Date.now()) {
const groups = new Map()
for (const a of abuse || []) {
const client = (clients || {})[a.client] || {}
const ip = client.ip || ''
const g = groups.get(ip) || {
ip,
pathCounts: new Map(),
clientHashes: new Set(),
lastStart: 0,
lastClient: a.client,
}
const start = new Date(a.start).getTime()
if (start > g.lastStart) {
g.lastStart = start
g.lastClient = a.client
}
const path = a.path || ''
const existing = g.pathCounts.get(path) || {
path,
count: 0,
firstStart: start,
flag: a.flag || false,
is_404: a.is_404 || false,
}
existing.count += 1
if (start < existing.firstStart) existing.firstStart = start
if (a.flag) existing.flag = true
if (!a.is_404) existing.is_404 = false
g.pathCounts.set(path, existing)
g.clientHashes.add(a.client)
groups.set(ip, g)
}
const totalHits = (g) => {
let n = 0
for (const p of g.pathCounts.values()) n += p.count
return n
}
return [...groups.values()]
.sort((a, b) => b.lastStart - a.lastStart)
.slice(0, 10)
.map((g) => {
const pathCategory = (p) => (p.flag ? 0 : p.is_404 ? 1 : 2)
const paths = [...g.pathCounts.values()].sort(
(a, b) =>
pathCategory(a) - pathCategory(b) ||
b.count - a.count ||
a.firstStart - b.firstStart,
)
const client = (clients || {})[g.lastClient] || {}
const host = client.host || ''
const isHost = !!host
return {
lastSeen: formatWhen(g.lastStart, now),
lastSeenIso: formatWhenIso(g.lastStart),
lastSeenLocal: formatWhenLocal(g.lastStart),
paths: paths.map((p) => ({
path: p.path,
count: p.count,
flag: p.flag,
is_404: p.is_404,
})),
allPaths: paths
.map((p) => (p.count > 1 ? `${p.count}× ${p.path}` : p.path))
.join('\n'),
clientCount: g.clientHashes.size,
ip: client.ip || g.ip,
ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip || g.ip) || client.ip || g.ip || '—',
isHost,
ua: client.ua_pretty || client.ua || '—',
uaRaw: client.ua || '',
lang: client.lang || '—',
langDisplay: formatLang(client.lang),
country: client.country || '—',
city: client.city || '—',
total: totalHits(g),
}
})
} }
/** /**
* Format raw visit records as rows for a technical table. Returns objects * Format raw visit records as rows for a technical table. Returns objects
* with display strings; missing values become "—". ``trail`` joins page * with display strings; missing values become "—". ``trail`` starts with the
* titles (when known) with " -> ". * external referer (when present), then the entry page and any further internal
* pages or external exit origins. Only the 20 most recent visits are shown.
* ``clients`` maps client hashes to client records.
*/ */
export function formatVisitRows(visits, pageTree) { export function formatVisitRows(visits, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree) const titles = buildTitleMap(pageTree)
return [...(visits || [])].reverse().map((v) => { return [...(visits || [])].reverse().slice(0, 20).map((v) => {
const client = (clients || {})[v.client] || {}
const trail = [v.entry, ...(v.trail || [])] const trail = [v.entry, ...(v.trail || [])]
.filter((p) => p?.startsWith('/')) .map((p) => stepOf(p, titles))
.map((p) => ({ .filter(Boolean)
path: p, const utmKeys = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content']
slug: slugOf(p), const utmValues = utmKeys.map((k) => (v.utm || {})[k]).filter(Boolean)
title: titles.get(p) || '', const utm = utmValues.length ? utmValues.join(' · ') : ''
})) const utmTitle = Object.entries(v.utm || {})
const utm = Object.entries(v.utm || {})
.map(([k, value]) => `${k}=${value}`) .map(([k, value]) => `${k}=${value}`)
.join(', ') .join(', ')
const dash = (s) => (s || '—') const dash = (s) => (s || '—')
const host = client.host || ''
const isHost = !!host
return { return {
when: new Date(v.start).toLocaleString(), lastSeen: formatWhen(v.start, now),
lastSeenIso: formatWhenIso(v.start),
lastSeenLocal: formatWhenLocal(v.start),
langDisplay: formatLang(client.lang),
trail, trail,
refererStep: stepOf(v.referer, titles),
referer: dash(v.referer), referer: dash(v.referer),
ip: v.ip || '', ip: client.ip || '',
ipDisplay: v.host || hostIP(v.ip) || v.ip || '—', ipDisplay: isHost ? mainDomain(host) : hostIP(client.ip) || client.ip || '—',
host: dash(v.host), isHost,
lang: dash(v.lang), lang: dash(client.lang),
country: dash(v.country), country: dash(client.country),
ua: v.ua_pretty || v.ua || '—', city: dash(client.city),
uaRaw: v.ua || '', ua: client.ua_pretty || client.ua || '',
uaRaw: client.ua || '',
utm: utm || '—', utm: utm || '—',
utmTitle,
} }
}) })
} }
+36 -7
View File
@@ -13,10 +13,11 @@ export const DAY = 86400e3
export const WEEK = 7 * DAY export const WEEK = 7 * DAY
export const RANGES = { export const RANGES = {
day: { label: 'day', span: DAY, bucket: MIN5 },
week: { label: 'week' }, week: { label: 'week' },
month: { label: 'month', span: 30 * DAY, bucket: 6 * HOUR }, month: { label: 'month', span: 30 * DAY, bucket: 6 * HOUR },
year: { label: 'year', span: 365 * DAY, bucket: DAY }, year: { label: 'year', span: 365 * DAY, bucket: DAY },
all: { label: 'all', span: null, bucket: DAY }, all: { label: 'all', span: null, bucket: DAY, minSpan: 30 * DAY },
} }
/** Monday 00:00 UTC of the week containing t (epoch day 0 was a Thursday). */ /** Monday 00:00 UTC of the week containing t (epoch day 0 was a Thursday). */
@@ -89,16 +90,19 @@ export function weeklySeries(buckets) {
/** /**
* Rolling window for the non-week ranges (x max = now), counts converted * Rolling window for the non-week ranges (x max = now), counts converted
* to per-day rates (the unit the month+ charts are read in). * to per-day rates (the unit the month+ charts are read in).
* Ranges without a fixed span use the full data reach, but never less than
* their configured minSpan so the chart keeps a readable minimum x scale.
*/ */
export function rollingSeries(buckets, rangeKey) { export function rollingSeries(buckets, rangeKey) {
const raw = rawTimes(buckets) const raw = rawTimes(buckets)
const times = Object.keys(raw).map(Number) const times = Object.keys(raw).map(Number)
if (!times.length) return null if (!times.length) return null
const { span, bucket } = RANGES[rangeKey] const { span, bucket, minSpan = 0 } = RANGES[rangeKey]
const t1 = Math.floor(Date.now() / bucket) * bucket + bucket const t1 = Math.floor(Date.now() / bucket) * bucket + bucket
const earliest = Math.floor(Math.min(...times) / bucket) * bucket
const t0 = span != null const t0 = span != null
? t1 - span ? t1 - span
: Math.floor(Math.min(...times) / bucket) * bucket : Math.min(earliest, t1 - minSpan)
const points = [] const points = []
for (let t = t0; t < t1; t += bucket) { for (let t = t0; t < t1; t += bucket) {
points.push({ t, count: sumRange(raw, t, t + bucket) }) points.push({ t, count: sumRange(raw, t, t + bucket) })
@@ -114,11 +118,36 @@ export function rollingSeries(buckets, rangeKey) {
} }
} }
/** Dispatch to weekly or rolling series based on the selected range. */ /**
* Day view: raw 5-minute bucket counts for the current 24-hour window.
* No smoothing or rate conversion is applied; counts are used as-is.
*/
export function daySeries(buckets) {
const raw = rawTimes(buckets)
const now = Date.now()
const { span, bucket } = RANGES.day
const t1 = Math.floor(now / bucket) * bucket + bucket
const t0 = t1 - span
const points = []
for (let t = t0; t < t1; t += bucket) {
points.push({ t, count: raw[t] || 0 })
}
return {
series: [{ points, label: '', opacity: 1, area: false }],
t0,
t1,
rate: 1,
binMinutes: bucket / 60e3,
unitMinutes: bucket / 60e3,
unit: '5min',
}
}
/** Dispatch to daily, weekly or rolling series based on the selected range. */
export function makeSeries(buckets, rangeKey) { export function makeSeries(buckets, rangeKey) {
return rangeKey === 'week' if (rangeKey === 'day') return daySeries(buckets)
? weeklySeries(buckets) if (rangeKey === 'week') return weeklySeries(buckets)
: rollingSeries(buckets, rangeKey) return rollingSeries(buckets, rangeKey)
} }
/** /**
+278 -88
View File
@@ -13,31 +13,39 @@
* connections. Animated beads flow along every edge in each direction, * connections. Animated beads flow along every edge in each direction,
* emitted at time intervals inversely proportional (linear) to the * emitted at time intervals inversely proportional (linear) to the
* directional count. * directional count.
* External referers appear as nodes in a row above the map, external exits * External sources appear as nodes in a row above the map. Sources are
* as small nodes just outside their source page, angled away from the * identified from visit records in this order: utm_campaign, utm_source,
* center. Self-loops (reload pings) are skipped. * referer, then other utm_* tags. Visits with a UTM tag are grouped under
* that tag's value, not under the referer domain. A UTM source node only
* becomes a clickable link when every visit carrying that UTM tag came
* from the same referer. External exits are full-size nodes just outside
* their source page, angled away from the center. Each distinct full exit
* URL is its own node. Self-loops (reload pings) are skipped.
*/ */
import { MIN_READ_SECONDS } from './format.js'
export const TNODE_R = 34 // node circles hold the slug and the view count export const TNODE_R = 34 // node circles hold the slug and the view count
export const EXT_R = 16 // external referer/exit nodes export const EXT_R = 34 // external referer/exit nodes use the same full size
// Edge width (half-width of the thin middle) grows logarithmically with // Edge width (half-width of the thin middle) grows logarithmically with
// the count, anchored so a single recorded transition renders as a ~1 px // the count. The constants are scaled down by ~10× so busy ranges (day,
// line. There is no cap — growth is slow enough that even very hot // year) do not overwhelm the graph with fat connectors. A single recorded
// connections stay reasonable. Connections carrying less than // transition still renders as a faint ~0.4 px line. Connections carrying
// PRUNE_FRACTION of the total traffic are not drawn at all (this also // less than PRUNE_FRACTION of the total traffic are not drawn at all (this
// keeps the number of drawn connections under ~100). // also keeps the graph under ~100 connections).
const WMID_MIN = 0.5 const WMID_MIN = 0.2
const WIDTH_GROWTH = 1.5 const WIDTH_GROWTH = 0.15
const PRUNE_FRACTION = 0.01 const PRUNE_FRACTION = 0.01
// Beads: each edge direction emits beads at count * BEAD_RATE beads per // Beads: each edge direction emits beads at count * BEAD_RATE beads per
// second (linear in the count). The component simulates every bead // second (linear in the count). The rate is reduced ~10× across all time
// independently in JS at BEAD_SPEED along the edge, with no limit on // scales to keep the animation lightweight. The component simulates every
// bead independently in JS at BEAD_SPEED along the edge, with no limit on
// beads in flight. // beads in flight.
export const BEAD_SPEED = 180 // svg units per second export const BEAD_SPEED = 180 // svg units per second
export const BEAD_R = 2.2 export const BEAD_R = 2.2
const BEAD_RATE = 0.12 // beads per second per recorded transition const BEAD_RATE = 0.012 // beads per second per recorded transition
const FLOW_OFFSET = 3 // lane offset to the right of the travel direction const FLOW_OFFSET = 3 // lane offset to the right of the travel direction
const MAX_EXT_IN = 8 // referer nodes in the top row const MAX_EXT_IN = 8 // referer nodes in the top row
@@ -83,30 +91,31 @@ function collectInternalTransitions(transitions) {
return internal return internal
} }
/** Short display label for an external origin (protocol stripped). */ /** Domain-only label for an external origin (path and www. removed). */
function extLabel(ext) { function extLabel(ext) {
const s = ext.replace(/^https?:\/\//, '') try {
return s.length > 18 ? `${s.slice(0, 17)}` : s const host = new URL(ext).hostname.replace(/^www\./, '')
return host.length > 25 ? `${host.slice(0, 24)}` : host
} catch {
const s = ext.replace(/^https?:\/\//, '').replace(/^www\./, '').split('/')[0]
return s.length > 25 ? `${s.slice(0, 24)}` : s
}
} }
/** /**
* Collect external transitions: referer origin -> entry page (incoming) and * Collect outgoing external transitions: page path -> full exit URL.
* page -> exit origin (outgoing). Aggregated per (origin, page) pair, with * Aggregated per (URL, page) pair. Incoming external links are now derived
* separate directional counts. "(direct)" entries are not links and skipped. * from visit records (which carry UTM tags), so only exits remain here.
*/ */
function collectExternalPairs(transitions) { function collectExitPairs(transitions) {
const pairs = new Map() // `${ext} ${page}` -> {ext, page, in, out} const pairs = new Map() // `${ext} ${page}` -> {ext, page, out}
for (const [fr, tos] of Object.entries(transitions || {})) { for (const [fr, tos] of Object.entries(transitions || {})) {
if (!fr.startsWith('/')) continue // ignore external -> anything
for (const [to, count] of Object.entries(tos)) { for (const [to, count] of Object.entries(tos)) {
const frExt = !fr.startsWith('/') if (!to.startsWith('http')) continue
const toExt = !to.startsWith('/') const k = `${to} ${fr}`
if (frExt === toExt) continue // internal-internal or ext-ext const p = pairs.get(k) || { ext: to, page: fr, out: 0 }
const ext = frExt ? fr : to p.out += count
const page = frExt ? to : fr
if (!ext.startsWith('http')) continue
const k = `${ext} ${page}`
const p = pairs.get(k) || { ext, page, in: 0, out: 0 }
p[frExt ? 'in' : 'out'] += count
pairs.set(k, p) pairs.set(k, p)
} }
} }
@@ -174,8 +183,29 @@ function layoutAngles(root, unit, weight) {
} }
} }
/** Compute median reading time per article in whole minutes. */
function buildReadMinutes(visits) {
const times = {}
for (const v of visits || []) {
for (const [path, sec] of Object.entries(v.read || {})) {
if (sec >= MIN_READ_SECONDS) {
;(times[path] || (times[path] = [])).push(sec)
}
}
}
const minutes = {}
for (const [path, arr] of Object.entries(times)) {
arr.sort((a, b) => a - b)
const mid = Math.floor(arr.length / 2)
const median =
arr.length % 2 ? arr[mid] : (arr[mid - 1] + arr[mid]) / 2
minutes[path] = Math.max(1, Math.round(median / 60))
}
return minutes
}
/** Compute radial positions, view counts and labels for each node. */ /** Compute radial positions, view counts and labels for each node. */
function positionNodes(nodes, maxDepth, unit, viewsData, titles) { function positionNodes(nodes, maxDepth, unit, viewsData, titles, readMinutes) {
// Constant radial gap between rings, equal to the arc spacing of nodes // Constant radial gap between rings, equal to the arc spacing of nodes
// along a ring: leaf arc = unit * GAP, so GAP scales up with `unit` on // along a ring: leaf arc = unit * GAP, so GAP scales up with `unit` on
// sparse trees (where closing the circle forces wider arcs) and with // sparse trees (where closing the circle forces wider arcs) and with
@@ -195,10 +225,14 @@ function positionNodes(nodes, maxDepth, unit, viewsData, titles) {
n.x = Math.cos(n.angle) * r n.x = Math.cos(n.angle) * r
n.y = Math.sin(n.angle) * r n.y = Math.sin(n.angle) * r
n.views = viewCount(n.path) n.views = viewCount(n.path)
n.readMin = readMinutes[n.path] || 0
// Slug inside the circle; full title goes on the link title attribute. // Slug inside the circle; full title goes on the link title attribute.
const slug = n.path === '/' ? '🏠' : n.path.split('/').pop() const slug = n.path === '/' ? '🏠' : n.path.split('/').pop()
n.label = slug.length > 11 ? `${slug.slice(0, 10)}` : slug n.label = slug.length > 16 ? `${slug.slice(0, 15)}` : slug
n.title = titles.get(n.path) || '' n.title = titles.get(n.path) || ''
// Category (non-leaf) pages with no views in this window are left
// blank to keep the layout, but their circle/label is not drawn.
n.hidden = n.children.length > 0 && n.views === 0
} }
return { radius, GAP } return { radius, GAP }
@@ -231,11 +265,33 @@ function buildFamilyArcs(nodes, radius) {
arcs.push({ arcs.push({
d: `M ${Math.cos(a0) * r} ${Math.sin(a0) * r} ` d: `M ${Math.cos(a0) * r} ${Math.sin(a0) * r} `
+ `A ${r} ${r} 0 ${large} 1 ${Math.cos(a1) * r} ${Math.sin(a1) * r}`, + `A ${r} ${r} 0 ${large} 1 ${Math.cos(a1) * r} ${Math.sin(a1) * r}`,
r,
a0,
a1,
}) })
} }
return arcs return arcs
} }
/** Bounding box of a circular arc centred at the origin, sampled. */
function arcBounds(r, a0, a1) {
let x0 = Infinity
let y0 = Infinity
let x1 = -Infinity
let y1 = -Infinity
const steps = 36
for (let i = 0; i <= steps; i++) {
const t = a0 + (a1 - a0) * (i / steps)
const x = Math.cos(t) * r
const y = Math.sin(t) * r
if (x < x0) x0 = x
if (y < y0) y0 = y
if (x > x1) x1 = x
if (y > y1) y1 = y
}
return { x0, y0, x1, y1 }
}
/** Collapse opposite transition directions into one unordered pair per page pair. */ /** Collapse opposite transition directions into one unordered pair per page pair. */
function aggregatePairs(internal) { function aggregatePairs(internal) {
const pairs = new Map() // unordered pair key -> [countAB, countBA] const pairs = new Map() // unordered pair key -> [countAB, countBA]
@@ -256,7 +312,7 @@ const fmtPt = (p) => `${p[0].toFixed(2)} ${p[1].toFixed(2)}`
* `wMid` is the half-width of the thin middle (already strength-scaled by * `wMid` is the half-width of the thin middle (already strength-scaled by
* the caller); `ra`/`rb` are the radii of the node circles each end wraps. * the caller); `ra`/`rb` are the radii of the node circles each end wraps.
*/ */
function buildRibbon(a, b, ab, ba, wMid, ra = TNODE_R, rb = TNODE_R) { function buildRibbon(a, b, ab, ba, wMid, ra = TNODE_R, rb = TNODE_R, external = false) {
const count = ab + ba const count = ab + ba
const len = Math.hypot(b.x - a.x, b.y - a.y) || 1 const len = Math.hypot(b.x - a.x, b.y - a.y) || 1
const ux = (b.x - a.x) / len const ux = (b.x - a.x) / len
@@ -364,6 +420,7 @@ function buildRibbon(a, b, ab, ba, wMid, ra = TNODE_R, rb = TNODE_R) {
return { return {
d, d,
title: `${a.path}${b.path}: ${count} (${ab} / ${ba})`, title: `${a.path}${b.path}: ${count} (${ab} / ${ba})`,
external,
} }
} }
@@ -378,7 +435,7 @@ function buildRibbon(a, b, ab, ba, wMid, ra = TNODE_R, rb = TNODE_R) {
* edge run on parallel lanes instead of colliding. The component turns * edge run on parallel lanes instead of colliding. The component turns
* these into independently simulated beads. * these into independently simulated beads.
*/ */
function buildFlows(a, b, ra, rb, ab, ba) { function buildFlows(a, b, ra, rb, ab, ba, visualScale = 1) {
const len = Math.hypot(b.x - a.x, b.y - a.y) || 1 const len = Math.hypot(b.x - a.x, b.y - a.y) || 1
const ux = (b.x - a.x) / len const ux = (b.x - a.x) / len
const uy = (b.y - a.y) / len const uy = (b.y - a.y) / len
@@ -399,7 +456,7 @@ function buildFlows(a, b, ra, rb, ab, ba) {
x2: a.x + toT * ux + s * rx, x2: a.x + toT * ux + s * rx,
y2: a.y + toT * uy + s * ry, y2: a.y + toT * uy + s * ry,
len: span, len: span,
interval: 1 / (count * BEAD_RATE), interval: 1 / (count * BEAD_RATE * visualScale),
} }
} }
const flows = [] const flows = []
@@ -414,14 +471,17 @@ function buildFlows(a, b, ra, rb, ab, ba) {
* Absolute on purpose — cool routes stay visible regardless of how hot * Absolute on purpose — cool routes stay visible regardless of how hot
* the hottest connection is. * the hottest connection is.
*/ */
const scaledWidth = (count) => WMID_MIN + WIDTH_GROWTH * Math.log1p(count - 1) const scaledWidth = (count) => {
if (count <= 0) return 0
return WMID_MIN + WIDTH_GROWTH * Math.log1p(count - 1)
}
/** /**
* Build ribbon edges and bead flows for every aggregated page-to-page * Build ribbon edges and bead flows for every aggregated page-to-page
* pair. Pairs carrying less than PRUNE_FRACTION of the total internal * pair. Pairs carrying less than PRUNE_FRACTION of the total internal
* traffic are pruned (this naturally bounds the graph to ~100 edges). * traffic are pruned (this naturally bounds the graph to ~100 edges).
*/ */
function buildInternalEdges(pairs, byPath) { function buildInternalEdges(pairs, byPath, visualScale = 1) {
let total = 0 let total = 0
for (const [, [ab, ba]] of pairs) total += ab + ba for (const [, [ab, ba]] of pairs) total += ab + ba
const minCount = total * PRUNE_FRACTION const minCount = total * PRUNE_FRACTION
@@ -433,8 +493,10 @@ function buildInternalEdges(pairs, byPath) {
const [pf, pt] = k.split(' ') const [pf, pt] = k.split(' ')
const a = byPath.get(pf) const a = byPath.get(pf)
const b = byPath.get(pt) const b = byPath.get(pt)
edges.push(buildRibbon(a, b, ab, ba, scaledWidth(ab + ba))) const wMid = scaledWidth((ab + ba) * visualScale)
flows.push(...buildFlows(a, b, TNODE_R, TNODE_R, ab, ba)) if (wMid <= 0) continue
edges.push(buildRibbon(a, b, ab, ba, wMid))
flows.push(...buildFlows(a, b, TNODE_R, TNODE_R, ab, ba, visualScale))
} }
return { edges, flows } return { edges, flows }
} }
@@ -475,42 +537,120 @@ export function filterViewsByRange(views, t0, t1) {
return filtered return filtered
} }
/** Keep only visits whose start time falls inside [t0, t1). */
export function filterVisitsByRange(visits, t0, t1) {
const out = []
for (const v of visits || []) {
const t = Date.parse(v.start)
if ((t0 == null || t >= t0) && (t1 == null || t < t1)) out.push(v)
}
return out
}
const UTM_PRIORITY = ['utm_campaign', 'utm_source']
const UTM_FALLBACK = ['utm_medium', 'utm_content', 'utm_term', 'utm_id']
/** Identify the source of a visit according to the requested priority. */
function identifySource(visit) {
const utm = visit.utm || {}
for (const k of UTM_PRIORITY) {
const v = utm[k]
if (v) return { value: v, isUtm: true }
}
if (visit.referer?.startsWith('http')) {
return { value: visit.referer, isUtm: false }
}
for (const k of UTM_FALLBACK) {
const v = utm[k]
if (v) return { value: v, isUtm: true }
}
return null
}
/** /**
* Place external referer and exit nodes and build their edges and bead * Collect source -> entry page pairs from visit records. Sources are
* identified by UTM campaign/source (then referer, then other UTM tags).
* A UTM source only gets a link href when every visit using that source
* came from the same referer; referer sources always link to their origin.
*/
function collectSourcePairs(visits) {
const groups = new Map() // `${source}\0${page}` -> pair
for (const v of visits || []) {
const src = identifySource(v)
if (!src) continue
const k = `${src.value}\0${v.entry}`
const p = groups.get(k) || {
source: src.value,
page: v.entry,
in: 0,
refs: new Set(),
missingRef: false,
href: null,
isUtm: src.isUtm,
}
p.in += 1
if (v.referer?.startsWith('http')) {
p.refs.add(v.referer)
} else {
p.missingRef = true
}
groups.set(k, p)
}
for (const p of groups.values()) {
if (p.isUtm && !p.missingRef && p.refs.size === 1) {
const ref = [...p.refs][0]
if (ref.startsWith('http')) p.href = ref
} else if (!p.isUtm && p.source.startsWith('http')) {
p.href = p.source
}
}
return [...groups.values()]
}
/**
* Place external source and exit nodes and build their edges and bead
* flows. * flows.
* Referers (incoming links) form a row centered above the map, hottest * Sources (incoming links) are derived from visit UTM/referer data and form
* first; exits sit just outside their source page, fanned away from the * a row centered above the map, hottest first; exits come from the
* center and nudged outwards until they no longer overlap any node. * transition matrix and sit just outside their source page.
* Widths and pruning use the same log scale and traffic-share rule as * Widths and pruning use the same log scale and traffic-share rule as
* internal connections. * internal connections.
*/ */
function buildExternal(external, byPath, radius, innerBounds) { function buildExternal({ sources, exits }, byPath, radius, innerBounds, visualScale = 1) {
const extNodes = [] const extNodes = []
const edges = [] const edges = []
const flows = [] const flows = []
let extTotal = 0 let extTotal = 0
for (const p of external) extTotal += p.in + p.out for (const p of sources) extTotal += p.in
for (const p of exits) extTotal += p.out
const minCount = extTotal * PRUNE_FRACTION const minCount = extTotal * PRUNE_FRACTION
const live = external.filter((p) => byPath.has(p.page)) const liveSources = sources.filter((p) => byPath.has(p.page))
if (!live.length) return { extNodes, edges, flows } const liveExits = exits.filter((p) => byPath.has(p.page))
if (!liveSources.length && !liveExits.length) return { extNodes, edges, flows }
const width = scaledWidth const width = (count) => scaledWidth(count * visualScale)
const overlaps = (x, y, r) => const overlaps = (x, y, r) =>
[...byPath.values(), ...extNodes].some( [...byPath.values(), ...extNodes].some(
(n) => Math.hypot(n.x - x, n.y - y) < (n.r ?? TNODE_R) + r + 10, (n) => Math.hypot(n.x - x, n.y - y) < (n.r ?? TNODE_R) + r + 10,
) )
// Incoming: one referer node per origin, in a row centered above the // Incoming: one source node per identified source, in a row centered
// map, with an edge to each page that origin led to. // above the map, with an edge to each page that source led to.
const byExt = new Map() // ext -> pairs, sorted by total incoming count const bySource = new Map() // source -> pairs, sorted by total incoming count
for (const p of live.filter((p) => p.in >= minCount)) { for (const p of liveSources.filter((p) => p.in >= minCount)) {
const g = byExt.get(p.ext) || [] const g = bySource.get(p.source) || []
g.push(p) g.push(p)
byExt.set(p.ext, g) bySource.set(p.source, g)
} }
const origins = [...byExt] const origins = [...bySource]
.map(([ext, ps]) => ({ ext, ps, total: ps.reduce((s, p) => s + p.in, 0) })) .map(([source, ps]) => ({
source,
ps,
total: ps.reduce((s, p) => s + p.in, 0),
href: ps[0].href,
isUtm: ps[0].isUtm,
}))
.sort((a, b) => b.total - a.total) .sort((a, b) => b.total - a.total)
.slice(0, MAX_EXT_IN) .slice(0, MAX_EXT_IN)
if (origins.length) { if (origins.length) {
@@ -518,43 +658,82 @@ function buildExternal(external, byPath, radius, innerBounds) {
const y = innerBounds.y0 - TNODE_R - 64 const y = innerBounds.y0 - TNODE_R - 64
const spacing = 2 * EXT_R + 44 const spacing = 2 * EXT_R + 44
const x0 = cx - ((origins.length - 1) * spacing) / 2 const x0 = cx - ((origins.length - 1) * spacing) / 2
origins.forEach(({ ext, ps }, i) => { origins.forEach(({ source, ps, total, href, isUtm }, i) => {
const xn = { path: ext, label: extLabel(ext), x: x0 + i * spacing, y, r: EXT_R } const label = isUtm ? source : extLabel(source)
const xn = {
path: source,
href,
label: label.length > 25 ? `${label.slice(0, 24)}` : label,
x: x0 + i * spacing,
y,
r: EXT_R,
count: total,
kind: 'source',
}
extNodes.push(xn) extNodes.push(xn)
for (const p of ps) { for (const p of ps) {
const page = byPath.get(p.page) const page = byPath.get(p.page)
edges.push(buildRibbon(xn, page, p.in, 0, width(p.in), EXT_R, TNODE_R)) const wMid = width(p.in)
flows.push(...buildFlows(xn, page, EXT_R, TNODE_R, p.in, 0)) if (wMid <= 0) continue
edges.push(buildRibbon(xn, page, p.in, 0, wMid, EXT_R, TNODE_R, true))
flows.push(...buildFlows(xn, page, EXT_R, TNODE_R, p.in, 0, visualScale))
} }
}) })
} }
// Outgoing: small exit nodes fanned outwards from the source page. // Outgoing: group by full URL so several links to the same domain stay
const outgoing = live.filter((p) => p.out >= minCount) // distinct. Each exit node is placed one ring-gap outside its source page
.sort((a, b) => b.out - a.out).slice(0, MAX_EXT_OUT) // (same radial spacing internal rings use), fanned around the source angle,
// and shows the total count across all pages that link to that URL.
const GAP = radius(1) - radius(0)
const outgoing = liveExits.filter((p) => p.out >= minCount)
.sort((a, b) => b.out - a.out)
const perPage = new Map() const perPage = new Map()
const selected = []
for (const p of outgoing) { for (const p of outgoing) {
const used = perPage.get(p.page) || 0 const used = perPage.get(p.page) || 0
if (used >= MAX_EXT_OUT_PER_PAGE) continue if (used >= MAX_EXT_OUT_PER_PAGE) continue
perPage.set(p.page, used + 1) perPage.set(p.page, used + 1)
selected.push(p)
if (selected.length >= MAX_EXT_OUT) break
}
const exitNodes = new Map() // full URL -> node
const placedPerPage = new Map() // for angle fanning of the placement anchor
for (const p of selected) {
const page = byPath.get(p.page) const page = byPath.get(p.page)
// Fan multiple exits of one page symmetrically around the outward let xn = exitNodes.get(p.ext)
// direction; the center page has no angle, so its exits point down if (!xn) {
// (the top row above the map belongs to referers). const used = placedPerPage.get(p.page) || 0
const base = page.depth ? page.angle : Math.PI / 2 placedPerPage.set(p.page, used + 1)
const ang = base + [0, 0.4, -0.4][used] const base = page.depth ? page.angle : Math.PI / 2
let dist = TNODE_R + 40 const ang = base + [0, 0.4, -0.4][used]
let x = page.x + Math.cos(ang) * dist let dist = GAP
let y = page.y + Math.sin(ang) * dist let x = page.x + Math.cos(ang) * dist
for (let tries = 0; tries < 5 && overlaps(x, y, EXT_R); tries++) { let y = page.y + Math.sin(ang) * dist
dist += 24 for (let tries = 0; tries < 5 && overlaps(x, y, EXT_R); tries++) {
x = page.x + Math.cos(ang) * dist dist += GAP * 0.3
y = page.y + Math.sin(ang) * dist x = page.x + Math.cos(ang) * dist
y = page.y + Math.sin(ang) * dist
}
xn = {
path: p.ext,
href: p.ext,
label: extLabel(p.ext),
x,
y,
r: EXT_R,
count: 0,
kind: 'exit',
}
exitNodes.set(p.ext, xn)
extNodes.push(xn)
} }
const xn = { path: p.ext, label: extLabel(p.ext), x, y, r: EXT_R } xn.count += p.out
extNodes.push(xn) const wMid = width(p.out)
edges.push(buildRibbon(page, xn, p.out, 0, width(p.out), TNODE_R, EXT_R)) if (wMid <= 0) continue
flows.push(...buildFlows(page, xn, TNODE_R, EXT_R, p.out, 0)) edges.push(buildRibbon(page, xn, p.out, 0, wMid, TNODE_R, EXT_R, true))
flows.push(...buildFlows(page, xn, TNODE_R, EXT_R, p.out, 0, visualScale))
} }
return { extNodes, edges, flows } return { extNodes, edges, flows }
@@ -565,11 +744,13 @@ function buildExternal(external, byPath, radius, innerBounds) {
* Returns { nodes, edges, flows, extNodes, arcs, bounds } or null when * Returns { nodes, edges, flows, extNodes, arcs, bounds } or null when
* there is nothing to show. * there is nothing to show.
*/ */
export function buildTransitionGraph(data, pageTree) { export function buildTransitionGraph(data, pageTree, visits = [], visualScale = 1) {
const internal = collectInternalTransitions(data?.transitions) const internal = collectInternalTransitions(data?.transitions)
const external = collectExternalPairs(data?.transitions) const sources = collectSourcePairs(visits)
const exits = collectExitPairs(data?.transitions)
const navOrder = buildNavigationOrder(pageTree) const navOrder = buildNavigationOrder(pageTree)
const titles = buildTitleMap(pageTree) const titles = buildTitleMap(pageTree)
const readMinutes = buildReadMinutes(visits)
if (!internal.length && !navOrder.size) return null if (!internal.length && !navOrder.size) return null
@@ -579,13 +760,14 @@ export function buildTransitionGraph(data, pageTree) {
layoutAngles(root, unit, weightFn) layoutAngles(root, unit, weightFn)
const maxDepth = Math.max(1, ...nodes.map((n) => n.depth)) const maxDepth = Math.max(1, ...nodes.map((n) => n.depth))
const { radius } = positionNodes(nodes, maxDepth, unit, data?.views, titles) const { radius } = positionNodes(nodes, maxDepth, unit, data?.views, titles, readMinutes)
const arcs = buildFamilyArcs(nodes, radius) const arcs = buildFamilyArcs(nodes, radius)
const pairs = aggregatePairs(internal) const pairs = aggregatePairs(internal)
const { edges, flows } = buildInternalEdges(pairs, byPath) const { edges, flows } = buildInternalEdges(pairs, byPath, visualScale)
// Tight bounding box of the actual page nodes; internal edges and arcs // Tight bounding box of the actual page nodes; family ring arcs can sweep
// stay within the node circles, so node bounds plus radius suffice. // outside the node circle (e.g. a large arc between two siblings on the
// left side reaching around the right), so their geometry is included too.
// External nodes extend the box below. // External nodes extend the box below.
const pad = 16 const pad = 16
const xs = nodes.map((n) => n.x) const xs = nodes.map((n) => n.x)
@@ -596,8 +778,16 @@ export function buildTransitionGraph(data, pageTree) {
x1: Math.max(...xs) + TNODE_R + pad, x1: Math.max(...xs) + TNODE_R + pad,
y1: Math.max(...ys) + TNODE_R + pad, y1: Math.max(...ys) + TNODE_R + pad,
} }
for (const arc of arcs) {
if (arc.a0 == null) continue
const b = arcBounds(arc.r, arc.a0, arc.a1)
bounds.x0 = Math.min(bounds.x0, b.x0)
bounds.y0 = Math.min(bounds.y0, b.y0)
bounds.x1 = Math.max(bounds.x1, b.x1)
bounds.y1 = Math.max(bounds.y1, b.y1)
}
const ext = buildExternal(external, byPath, radius, bounds) const ext = buildExternal({ sources, exits }, byPath, radius, bounds, visualScale)
for (const xn of ext.extNodes) { for (const xn of ext.extNodes) {
bounds.x0 = Math.min(bounds.x0, xn.x - xn.r - pad) bounds.x0 = Math.min(bounds.x0, xn.x - xn.r - pad)
bounds.y0 = Math.min(bounds.y0, xn.y - xn.r - pad) bounds.y0 = Math.min(bounds.y0, xn.y - xn.r - pad)
+111 -19
View File
@@ -338,29 +338,107 @@ import "overlayscrollbars/overlayscrollbars.css";
} }
// --- Analytics pings --------------------------------------------------- // --- Analytics pings ---------------------------------------------------
// Fire-and-forget POST /_a {fr, to}: on the initial page load (starts the // Fire-and-forget POST /_a {fr, to, read}: on the initial page load
// visit — the server counts nothing from the document GET alone), for // (starts the visit — the server counts nothing from the document GET
// internal fetch-navigations and for external https exits. Excluded: // alone), for internal fetch-navigations, for external https exits, and
// back/forward (popstate never pings) and everything while we know the // on window close. ``read`` is the active time (ms) spent on ``fr``.
// user is an admin — but only when SSO is actually in use; with no auth // Reading time pauses after 1 minute of inactivity and resumes on the
// (dev/test) "admin" is everyone's state and nothing would be recorded — // next mouse/touch/scroll/keyboard event.
// or has the editor open (admin noise, not visits). The analytics page // Excluded: back/forward (popstate never pings), everything while the
// itself (/_a) is also excluded even though fetch-navigation treats it like // editor is open (body.editing — admin noise, not visits), and the
// a normal article. // analytics page itself (/_a), even though fetch-navigation treats it
// like a normal article.
// Admins (when SSO is actually in use — with no auth proxy "admin" is
// everyone's state) ping normally but with hide=1: the server then
// records nothing and scrubs any session the same browser accumulated
// before logging in, so admins never show up as visits or crawlers.
// See docs/analytics.md. // See docs/analytics.md.
function ping(to, fr = currentPath) { function ping(to, fr = currentPath, read = 0) {
if ((ssoAvailable && isAdmin) || document.body.classList.contains("editing") if (document.body.classList.contains("editing")) return;
|| to === "/_a" || fr === "/_a") return; if ((to && to === "/_a") || fr === "/_a") return;
const hide = ssoAvailable && isAdmin ? 1 : 0;
const body = JSON.stringify({
fr, to, hide,
read: Math.max(0, Math.round(read / 1000)),
});
try { try {
fetch("/_a", { fetch("/_a", {
method: "POST", method: "POST",
keepalive: true, keepalive: true,
headers: { "content-type": "application/json" }, headers: { "content-type": "application/json" },
body: JSON.stringify({ fr, to }), body,
}); });
} catch { /* analytics must never break navigation */ } } catch { /* analytics must never break navigation */ }
} }
// Active reading time for the current page. The clock stops after 1 minute
// without activity and restarts on the next mouse/touch/scroll/keyboard
// event.
const INACTIVE_MS = 60_000;
let readStart = performance.now();
let readElapsed = 0;
let reading = true;
let readInactivityTimer = null;
let closePingedFor = null;
function markReadActivity() {
if (!reading) {
reading = true;
readStart = performance.now();
}
clearTimeout(readInactivityTimer);
readInactivityTimer = setTimeout(() => {
if (reading) {
readElapsed += performance.now() - readStart;
reading = false;
}
}, INACTIVE_MS);
}
function takeReadTime() {
if (reading) {
readElapsed += performance.now() - readStart;
readStart = performance.now();
}
const ms = Math.max(0, Math.round(readElapsed));
readElapsed = 0;
return ms;
}
function resetReadTime() {
readElapsed = 0;
reading = true;
readStart = performance.now();
clearTimeout(readInactivityTimer);
}
function sendClosePing() {
if (closePingedFor === currentPath) return;
const read = Math.max(0, Math.round(takeReadTime() / 1000));
if (read <= 0) return;
const hide = ssoAvailable && isAdmin ? 1 : 0;
const body = JSON.stringify({ fr: currentPath, hide, read });
const blob = new Blob([body], { type: "application/json" });
try {
if (navigator.sendBeacon) {
navigator.sendBeacon("/_a", blob);
} else {
fetch("/_a", {
method: "POST",
keepalive: true,
headers: { "content-type": "application/json" },
body,
});
}
} catch { /* analytics must never break navigation */ }
closePingedFor = currentPath;
}
for (const ev of ["mousemove", "mousedown", "touchstart", "touchmove", "scroll", "keydown"]) {
addEventListener(ev, markReadActivity, { passive: true });
}
addEventListener("pagehide", sendClosePing);
// The initial page load pings too — it is what starts the visit and // The initial page load pings too — it is what starts the visit and
// counts the entry page view (the document GET alone records nothing). // counts the entry page view (the document GET alone records nothing).
// Sent once per load, after the auth probes so the admin gate applies; // Sent once per load, after the auth probes so the admin gate applies;
@@ -425,7 +503,11 @@ import "overlayscrollbars/overlayscrollbars.css";
if (!res.ok || !type.includes("text/html")) throw new Error("not a page"); if (!res.ok || !type.includes("text/html")) throw new Error("not a page");
// Reflect any redirect the server issued. // Reflect any redirect the server issued.
if (res.redirected) finalUrl = res.url; if (res.redirected) finalUrl = res.url;
doc = new DOMParser().parseFromString(await res.text(), "text/html"); const html = await res.text();
// Populate the cache too, or the post-swap preload (which includes
// location.pathname) would fetch the very page we just loaded again.
pageCache.set(new URL(finalUrl, location.href).pathname, html);
doc = new DOMParser().parseFromString(html, "text/html");
} catch { } catch {
location.href = url; // fall back to a normal navigation location.href = url; // fall back to a normal navigation
return false; return false;
@@ -471,7 +553,9 @@ import "overlayscrollbars/overlayscrollbars.css";
runScripts(document.getElementById("page-banner")); runScripts(document.getElementById("page-banner"));
runScripts(document.getElementById("main")); runScripts(document.getElementById("main"));
applyEffects(); applyEffects();
mountAnalytics(document); // The fetched doc carries the analytics meta; the live document's
// <head> is never swapped, so querying it would never find the entry.
mountAnalytics(doc);
}; };
// Rotating cube page transition (see the FRAGILE block in pagerite.css); // Rotating cube page transition (see the FRAGILE block in pagerite.css);
// mirrored when navigating back through history. Navigation within the // mirrored when navigating back through history. Navigation within the
@@ -530,9 +614,12 @@ import "overlayscrollbars/overlayscrollbars.css";
if (!a || a.target || a.hasAttribute("download")) return; if (!a || a.target || a.hasAttribute("download")) return;
const url = new URL(a.href, location.href); const url = new URL(a.href, location.href);
if (url.origin !== location.origin) { if (url.origin !== location.origin) {
// External link: the browser navigates; just record the exit (https // External link: the browser navigates; record the full https URL so
// origins only, stripped to the origin part server-side anyway). // different links to the same domain stay distinct in analytics.
if (url.protocol === "https:") ping(url.origin); if (url.protocol === "https:") {
closePingedFor = currentPath;
ping(url.href, currentPath, takeReadTime());
}
return; return;
} }
// Same-page anchor links (footnotes etc.): let the browser handle them // Same-page anchor links (footnotes etc.): let the browser handle them
@@ -544,7 +631,12 @@ import "overlayscrollbars/overlayscrollbars.css";
ev.preventDefault(); ev.preventDefault();
// Capture the source now: load() updates currentPath before pinging. // Capture the source now: load() updates currentPath before pinging.
const from = currentPath; const from = currentPath;
load(url).then((ok) => { if (ok) ping(url.pathname, from); }); load(url).then((ok) => {
if (!ok) return;
closePingedFor = null;
ping(url.pathname, from, takeReadTime());
resetReadTime();
});
}); });
addEventListener("popstate", () => { addEventListener("popstate", () => {
+1 -1
View File
@@ -11,7 +11,7 @@
*/ */
export default function fastapiVue({ paths = ["/api"] } = {}) { export default function fastapiVue({ paths = ["/api"] } = {}) {
const backendUrl = process.env.PAGERITE_BACKEND_URL || "http://localhost:3200" const backendUrl = process.env.PAGERITE_BACKEND_URL || "http://localhost:8210"
// Build proxy configuration for each path // Build proxy configuration for each path
const proxy = {} const proxy = {}
+69 -2
View File
@@ -1,14 +1,74 @@
# auto-upgrade@fastapi-vue-setup - remove this if you modify this file
"""Command-line entry point for running the backend server.""" """Command-line entry point for running the backend server."""
import argparse import argparse
import gzip
import os import os
import sys
from datetime import date
from pathlib import Path
import httpx
from fastapi_vue import server from fastapi_vue import server
DEFAULT_PORT = 3100 DEFAULT_PORT = 8100
DEVMODE = os.getenv("PAGERITE_DEV") == "1" DEVMODE = os.getenv("PAGERITE_DEV") == "1"
# Repository root (pagerite/__main__.py -> ..), where the MMDB lives.
_REPO_ROOT = Path(__file__).resolve().parent.parent
DBIP_URL = "https://download.db-ip.com/free/dbip-city-lite-{month}.mmdb.gz"
def _download_dbip() -> None:
"""Download the latest dbip-city-lite MMDB if ours is missing or older."""
today = date.today()
months = [f"{today:%Y-%m}"]
# The current month's file may not be published yet; fall back to last month.
prev = (today.replace(day=1) - date.resolution).replace(day=1)
months.append(f"{prev:%Y-%m}")
existing = sorted(
p.stem.removeprefix("dbip-city-lite-").removesuffix(".mmdb")
for p in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*")
)
if existing and existing[-1] >= months[0]:
print(f"pagerite: DB-IP database is current ({existing[-1]}), skipping download")
return
for month in months:
url = DBIP_URL.format(month=month)
target = _REPO_ROOT / f"dbip-city-lite-{month}.mmdb.gz"
tmp = target.with_suffix(".mmdb.gz.tmp")
print(f"pagerite: downloading {url}")
try:
with httpx.stream("GET", url, follow_redirects=True, timeout=120) as r:
if r.status_code == 404:
continue
r.raise_for_status()
with open(tmp, "wb") as f:
for chunk in r.iter_bytes():
f.write(chunk)
except httpx.HTTPError as e:
print(f"pagerite: DB-IP download failed: {e}", file=sys.stderr)
tmp.unlink(missing_ok=True)
continue
# Verify it is actually gzip data before installing it.
try:
with gzip.open(tmp, "rb") as f:
f.read(1)
except OSError:
print(f"pagerite: DB-IP download for {month} was not valid gzip", file=sys.stderr)
tmp.unlink(missing_ok=True)
continue
os.replace(tmp, target)
# Drop older databases so the app never picks up a stale one.
for old in _REPO_ROOT.glob("dbip-city-lite-*.mmdb*"):
if old.name != target.name:
old.unlink()
print(f"pagerite: DB-IP database updated to {target.name}")
return
print("pagerite: could not download a DB-IP database", file=sys.stderr)
def main() -> None: def main() -> None:
"""Run the backend server with optional arguments.""" """Run the backend server with optional arguments."""
@@ -19,7 +79,14 @@ def main() -> None:
action="append", action="append",
help=(f"Endpoint (default: localhost:{DEFAULT_PORT})."), help=(f"Endpoint (default: localhost:{DEFAULT_PORT})."),
) )
parser.add_argument(
"--dbip",
action="store_true",
help="Download/update the DB-IP city lite database before starting.",
)
args = parser.parse_args() args = parser.parse_args()
if args.dbip:
_download_dbip()
dev = {"reload": True, "reload_dirs": ["pagerite"]} if DEVMODE else {} dev = {"reload": True, "reload_dirs": ["pagerite"]} if DEVMODE else {}
server.run( server.run(
"pagerite.app:app", "pagerite.app:app",
+441 -100
View File
@@ -5,22 +5,31 @@ ping on page load starts a visit, later pings extend it, and pings with no
known session start a fresh one (missing data, not dropped). The document known session start a fresh one (missing data, not dropped). The document
GET handler stashes the entry referer (external https origin) and any GET handler stashes the entry referer (external https origin) and any
utm_* query parameters in in-memory IP tables, consumed when the ping utm_* query parameters in in-memory IP tables, consumed when the ping
starts the visit; nothing is counted without a ping (bots and admin starts the visit; nothing is counted without a ping (bots stay invisible).
browsing stay invisible). The session map is in-memory only. The visitor Admin clients ping with ``hide=1``, which records nothing and removes any
IP and, when available, its reverse-DNS host name are stored on the visit visit the session accumulated before logging in. Scanner telltale 404s
record itself. (dotpaths, *.php) classify the source IP as abuse; its hits — including
earlier crawler hits — are moved to the abuse list, which the viewer
groups by IP with full request paths. Client metadata (IP, UA, language,
country/city, host) is stored once per unique client hash and referenced
from visits, crawler hits and abuse hits. The session map is in-memory
only.
Data is a msgspec Struct JSON-dumped to its own file (not the kanta db), Data is a msgspec Struct JSON-dumped to its own file (not the kanta db),
rewritten atomically on every recorded event. rewritten atomically on every recorded event.
""" """
import ipaddress
import os import os
import re import re
import tempfile import tempfile
from collections.abc import Callable
from contextlib import suppress
from datetime import UTC, datetime, timedelta from datetime import UTC, datetime, timedelta
from pathlib import Path from pathlib import Path
from urllib.parse import parse_qs, urlparse from urllib.parse import parse_qs, urlparse
import blake3
import msgspec import msgspec
from ua_parser import parse from ua_parser import parse
@@ -39,7 +48,10 @@ def _compact_user_agent(ua: str) -> str:
dev = r.device.family if r.device else None dev = r.device.family if r.device else None
if browser in (None, "Other") and os_name in (None, "Other"): if browser in (None, "Other") and os_name in (None, "Other"):
return ua return ua
browser = browser if browser and browser != "Other" else "" if browser and browser != "Other":
browser = browser.split()[0]
else:
browser = ""
os_name = os_name if os_name and os_name != "Other" else "" os_name = os_name if os_name and os_name != "Other" else ""
if dev in (None, "Other") or dev == browser: if dev in (None, "Other") or dev == browser:
dev = "" dev = ""
@@ -47,50 +59,90 @@ def _compact_user_agent(ua: str) -> str:
return " ".join(p for p in parts if p).strip() return " ".join(p for p in parts if p).strip()
class Client(msgspec.Struct, omit_defaults=True):
"""Client metadata shared by visits, crawler hits and abuse hits.
Identified by a 6-byte blake3 hash of the IPv4 address or IPv6 /64
network, the full User-Agent string and the extracted language tag.
Country/city/host are filled in asynchronously after the first event.
"""
#: Visitor IP address (first X-Forwarded-For hop or direct peer).
ip: str = ""
#: Reverse-DNS host name for ``ip`` when resolvable, else "".
host: str = ""
#: First Accept-Language tag, lowercased (e.g. "en-us").
lang: str = ""
#: Two-letter country code from the DB-IP geoip lookup, or "".
country: str = ""
#: City name from the DB-IP geoip lookup, or "".
city: str = ""
#: Raw User-Agent header.
ua: str = ""
#: Compact display form of ``ua`` (browser/OS/device) when parsable.
ua_pretty: str = ""
class Visit(msgspec.Struct, omit_defaults=True): class Visit(msgspec.Struct, omit_defaults=True):
"""One visit: the initial-load data plus everything seen afterwards. """One visit: the initial-load data plus everything seen afterwards.
``trail`` holds page paths and external exit origins in first-seen ``trail`` holds page paths and external exit URLs in first-seen
order; re-visiting an already seen page does not append. The entry order; re-visiting an already seen page does not append. The entry
page itself is in ``entry``, not in the trail. page itself is in ``entry``, not in the trail. Client metadata is
held in ``Analytics.clients`` keyed by ``client``.
""" """
start: datetime start: datetime
entry: str entry: str
#: External https origin of the initial load, "" for direct visits. #: External https origin of the initial load, "" for direct visits.
referer: str = "" referer: str = ""
#: Visitor IP address (first X-Forwarded-For hop or direct peer). #: 6-byte blake3 hash referencing ``Analytics.clients``.
ip: str = "" client: bytes = b""
#: Reverse-DNS host name for ``ip`` when resolvable, else "".
host: str = ""
trail: list[str] = [] trail: list[str] = []
#: First Accept-Language tag, lowercased (e.g. "en-us").
lang: str = ""
#: Two-letter region subtag derived from ``lang`` (e.g. "US"), or "".
country: str = ""
#: Raw User-Agent header from the initial ping.
ua: str = ""
#: Compact display form of ``ua`` (browser/OS/device) when parsable.
ua_pretty: str = ""
#: UTM query parameters from the landing URL, keyed by parameter name. #: UTM query parameters from the landing URL, keyed by parameter name.
utm: dict[str, str] = {} utm: dict[str, str] = {}
#: Active reading time per path (seconds), keyed by path.
read: dict[str, int] = {}
class CrawlerHit(msgspec.Struct, omit_defaults=True): class CrawlerHit(msgspec.Struct, omit_defaults=True):
"""A document GET that was never followed by an analytics ping.""" """A document GET that was never followed by an analytics ping.
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
"""
start: datetime start: datetime
entry: str entry: str
ip: str = "" #: 6-byte blake3 hash referencing ``Analytics.clients``.
ua: str = "" client: bytes = b""
#: Compact display form of ``ua`` when parsable.
ua_pretty: str = ""
#: External https origin of the initial load, "" for direct/none. #: External https origin of the initial load, "" for direct/none.
referer: str = "" referer: str = ""
#: Raw query string of the landing URL (UTM tags can be parsed from it). #: Raw query string of the landing URL (UTM tags can be parsed from it).
query: str = "" query: str = ""
class AbuseHit(msgspec.Struct, omit_defaults=True):
"""A request from an IP classified as a scanner/abuser.
Unlike crawler hits the full request path (query string included) is
kept: the interesting part is exactly which paths were probed.
``flag`` marks the path that triggered classification; ``is_404``
distinguishes 404 responses from document GETs made by the abuser.
Client metadata is held in ``Analytics.clients`` keyed by ``client``.
"""
start: datetime
#: Full request path including the query string (e.g. "/.env?x=1").
path: str
#: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b""
#: True when this path triggered abuse classification (telltale path
#: or the 404 that crossed the threshold).
flag: bool = False
#: True for 404 responses; false for document GETs from the abuser.
is_404: bool = False
class Analytics(msgspec.Struct, omit_defaults=True): class Analytics(msgspec.Struct, omit_defaults=True):
"""Root of the analytics JSON file. Append-only by design: old data is """Root of the analytics JSON file. Append-only by design: old data is
dropped by deleting list entries / bucket keys.""" dropped by deleting list entries / bucket keys."""
@@ -98,6 +150,12 @@ class Analytics(msgspec.Struct, omit_defaults=True):
visits: list[Visit] = [] visits: list[Visit] = []
#: Document GETs that never produced a ping, treated as crawler/bot hits. #: Document GETs that never produced a ping, treated as crawler/bot hits.
crawlers: list[CrawlerHit] = [] crawlers: list[CrawlerHit] = []
#: Requests from abusive IPs (see AbuseHit), grouped by IP in the viewer.
abuse: list[AbuseHit] = []
#: Client metadata keyed by 6-byte blake3 hash.
clients: dict[bytes, Client] = {}
#: IPs classified as scanners/abusers (keys; values always True).
abuse_ips: dict[str, bool] = {}
#: Page transitions per 5-minute bucket (sparse): #: Page transitions per 5-minute bucket (sparse):
#: from -> to -> bucket ISO -> count. ``from`` is the referer origin or #: from -> to -> bucket ISO -> count. ``from`` is the referer origin or
#: "(direct)" for initial loads, a page path for pings. #: "(direct)" for initial loads, a page path for pings.
@@ -124,6 +182,17 @@ def _origin(url: str) -> str | None:
return f"https://{parsed.netloc}" return f"https://{parsed.netloc}"
def _external_target(url: str) -> str | None:
"""A valid https URL (origin or full page), else None."""
try:
parsed = urlparse(url)
except ValueError:
return None
if parsed.scheme != "https" or not parsed.netloc:
return None
return url
_SEGMENT = re.compile(r"[a-z0-9][a-z0-9_-]*") _SEGMENT = re.compile(r"[a-z0-9][a-z0-9_-]*")
@@ -170,9 +239,50 @@ def _utm_tags(query: str) -> dict[str, str]:
_CRAWLER_TIMEOUT = timedelta(seconds=10) _CRAWLER_TIMEOUT = timedelta(seconds=10)
#: Plain-404 count per IP that classifies it as abuse even without a
#: telltale path hit.
_ABUSE_404_THRESHOLD = 10
#: Paths that instantly classify an IP as abuse when they 404: any segment
#: starting with a dot ("/.env", "/.git/config") or ending in ".php".
_ABUSE_PATH = re.compile(r"(^|/)\.|\.php$", re.IGNORECASE)
def _is_abuse_path(path: str) -> bool:
"""Telltale scanner path: dot segment or *.php."""
return bool(_ABUSE_PATH.search(path.split("?")[0]))
def _network_ip(ip: str) -> str:
"""IPv4 address unchanged, IPv6 collapsed to its /64 network address.
We hash the network rather than the full address so that clients in the
same /64 (a typical end-user allocation) are treated as one visitor.
"""
if not ip:
return ip
try:
addr = ipaddress.ip_address(ip)
except ValueError:
return ip
if isinstance(addr, ipaddress.IPv6Address):
return str(ipaddress.IPv6Network(f"{ip}/64", strict=False).network_address)
return ip
def _client_hash(ip: str, ua: str, lang: str) -> bytes:
"""6-byte blake3 digest identifying a visitor/client tuple.
The key is the prettified IP (IPv6 /64), the raw UA string and the
extracted language tag, separated by null bytes.
"""
return blake3.blake3(
f"{_network_ip(ip)}\0{ua}\0{lang}".encode()
).digest()[:6]
class Store: class Store:
"""In-memory analytics data plus the (IP, UA) -> visit session map.""" """In-memory analytics data plus the client-hash -> visit session map."""
def __init__(self, path: Path) -> None: def __init__(self, path: Path) -> None:
self.path = path self.path = path
@@ -182,8 +292,8 @@ class Store:
self.data = msgspec.json.decode(path.read_bytes(), type=Analytics) self.data = msgspec.json.decode(path.read_bytes(), type=Analytics)
except msgspec.DecodeError, OSError: except msgspec.DecodeError, OSError:
pass # legacy schema / corrupt or unreadable file: start fresh pass # legacy schema / corrupt or unreadable file: start fresh
#: (ip, user-agent) -> index of the current visit in data.visits #: client hash -> index of the current visit in data.visits
self.sessions: dict[tuple[str, str], int] = {} self.sessions: dict[bytes, int] = {}
#: ip -> external https origin of the latest document GET carrying #: ip -> external https origin of the latest document GET carrying
#: one, stashed for the visit the client's initial ping starts. #: one, stashed for the visit the client's initial ping starts.
#: Internal or absent referers never touch the table. #: Internal or absent referers never touch the table.
@@ -196,6 +306,26 @@ class Store:
#: Document GETs that have not yet been matched by a ping. Kept #: Document GETs that have not yet been matched by a ping. Kept
#: in RAM only; expired entries are written to ``data.crawlers``. #: in RAM only; expired entries are written to ``data.crawlers``.
self.pending_crawlers: list[CrawlerHit] = [] self.pending_crawlers: list[CrawlerHit] = []
#: ip -> number of plain (non-telltale) 404s seen, in RAM only;
#: reaching ``_ABUSE_404_THRESHOLD`` classifies the IP as abuse.
self.not_found_counts: dict[str, int] = {}
#: Callables to notify when persisted data changes. Registered by the
#: analytics WebSocket broadcaster.
self._on_change: list[Callable[[], None]] = []
def subscribe(self, callback: Callable[[], None]) -> None:
"""Register a callback to be called after every persisted change."""
if callback not in self._on_change:
self._on_change.append(callback)
def unsubscribe(self, callback: Callable[[], None]) -> None:
"""Remove a previously registered change callback."""
with suppress(ValueError):
self._on_change.remove(callback)
def _notify(self) -> None:
for callback in self._on_change:
callback()
def _save(self) -> None: def _save(self) -> None:
"""Rewrite the JSON file atomically (temp file + rename).""" """Rewrite the JSON file atomically (temp file + rename)."""
@@ -208,21 +338,29 @@ class Store:
os.replace(tmp, self.path) os.replace(tmp, self.path)
except OSError: except OSError:
pass # analytics must never break page serving pass # analytics must never break page serving
else:
self._notify()
def _flush_crawlers(self, now: datetime | None = None) -> None: def _flush_crawlers(self, now: datetime | None = None) -> list[bytes]:
"""Move expired pending crawler hits into persistent ``data.crawlers``.""" """Move expired pending crawler hits into persistent ``data.crawlers``.
Returns the client hashes of the newly flushed hits so callers can
schedule async enrichment.
"""
if not self.pending_crawlers: if not self.pending_crawlers:
return return []
now = now or datetime.now(UTC) now = now or datetime.now(UTC)
cutoff = now - _CRAWLER_TIMEOUT cutoff = now - _CRAWLER_TIMEOUT
expired: list[CrawlerHit] = [] expired: list[CrawlerHit] = []
remaining: list[CrawlerHit] = [] remaining: list[CrawlerHit] = []
for hit in self.pending_crawlers: for hit in self.pending_crawlers:
(expired if hit.start <= cutoff else remaining).append(hit) (expired if hit.start <= cutoff else remaining).append(hit)
if expired: if not expired:
self.pending_crawlers = remaining return []
self.data.crawlers.extend(expired) self.pending_crawlers = remaining
self._save() self.data.crawlers.extend(expired)
self._save()
return [hit.client for hit in expired]
def _count(self, table: dict[str, int], key: str) -> None: def _count(self, table: dict[str, int], key: str) -> None:
table[key] = table.get(key, 0) + 1 table[key] = table.get(key, 0) + 1
@@ -232,15 +370,188 @@ class Store:
buckets = self.data.transitions.setdefault(fr, {}).setdefault(to, {}) buckets = self.data.transitions.setdefault(fr, {}).setdefault(to, {})
self._count(buckets, _bucket(now)) self._count(buckets, _bucket(now))
def _uncount(self, table: dict[str, int], key: str) -> None:
"""Reverse one ``_count``: decrement and drop empty keys."""
if key in table:
table[key] -= 1
if table[key] <= 0:
del table[key]
def _remove_visit(self, index: int) -> None:
"""Delete a visit and reverse the counts its creation recorded.
Used when a known visitor turns out to be an admin (hide=1 ping):
the session is scrubbed from the stats. Views/transitions logged
by later pings inside the visit lack per-event timestamps and are
left as-is.
"""
visit = self.data.visits[index]
bucket = _bucket(visit.start)
self._uncount(self.data.site_visits, bucket)
views = self.data.views.get(visit.entry)
if views is not None:
self._uncount(views, bucket)
if not views:
del self.data.views[visit.entry]
fr_map = self.data.transitions.get(visit.referer or "(direct)")
if fr_map is not None:
buckets = fr_map.get(visit.entry)
if buckets is not None:
self._uncount(buckets, bucket)
if not buckets:
del fr_map[visit.entry]
if not fr_map:
del self.data.transitions[visit.referer or "(direct)"]
del self.data.visits[index]
# Sessions store list indices; shift the ones past the removed visit.
for key, i in list(self.sessions.items()):
if i > index:
self.sessions[key] = i - 1
def _client_ip(self, client_hash: bytes) -> str:
"""Return the IP stored for ``client_hash``, or "" if missing."""
client = self.data.clients.get(client_hash)
return client.ip if client else ""
def _ensure_client(
self,
ip: str,
ua: str,
lang: str,
*,
country: str = "",
) -> bytes:
"""Get or create a ``Client`` record; return its 6-byte hash."""
h = _client_hash(ip, ua, lang)
if h not in self.data.clients:
self.data.clients[h] = Client(
ip=ip,
ua=ua,
ua_pretty=_compact_user_agent(ua),
lang=lang,
country=country,
)
self._save()
return h
def enrich_client(
self,
client_hash: bytes,
*,
host: str = "",
country: str = "",
city: str = "",
) -> None:
"""Fill in host/geoip fields on a client record after async lookups."""
client = self.data.clients.get(client_hash)
if client is None:
return
changed = False
if host and not client.host:
client.host = host
changed = True
if country:
client.country = country
changed = True
if city:
client.city = city
changed = True
if changed:
self._save()
def _abuse_hit(
self,
client_hash: bytes,
path: str,
start: datetime | None = None,
*,
flag: bool = False,
is_404: bool = False,
) -> None:
"""Append one abuse hit referencing a client by hash."""
self.data.abuse.append(
AbuseHit(
start=start or datetime.now(UTC),
path=path,
client=client_hash,
flag=flag,
is_404=is_404,
)
)
def classify_abuse(
self,
ip: str,
client_hash: bytes,
path: str,
*,
flag: bool = False,
is_404: bool = False,
) -> None:
"""Classify an IP as a scanner/abuser and record the triggering hit.
All earlier crawler hits from the same IP (persisted and pending)
are moved to the abuse list — a random-UA scanner must not pollute
the crawler stats of the legitimate bots it impersonates.
"""
if ip not in self.data.abuse_ips:
self.data.abuse_ips[ip] = True
moved = [h for h in self.data.crawlers if self._client_ip(h.client) == ip]
if moved:
self.data.crawlers = [h for h in self.data.crawlers if self._client_ip(h.client) != ip]
for h in moved:
self._abuse_hit(
h.client,
h.entry + (f"?{h.query}" if h.query else ""),
start=h.start,
)
pending = [h for h in self.pending_crawlers if self._client_ip(h.client) == ip]
if pending:
self.pending_crawlers = [h for h in self.pending_crawlers if self._client_ip(h.client) != ip]
for h in pending:
self._abuse_hit(
h.client,
h.entry + (f"?{h.query}" if h.query else ""),
start=h.start,
)
self._abuse_hit(client_hash, path, flag=flag, is_404=is_404)
self._save()
def track_404(
self,
ip: str,
ua: str,
path: str,
accept_language: str = "",
) -> bytes:
"""Record a 404 response for ``path`` (full path, query included).
A telltale path (dot segment or *.php) classifies the IP as abuse
immediately; enough plain 404s from one IP do too. Hits from
already-classified IPs go straight to the abuse list.
Returns the client hash so callers can schedule async enrichment.
"""
lang, country = _parse_accept_language(accept_language)
client_hash = self._ensure_client(ip, ua, lang, country=country)
if ip in self.data.abuse_ips:
self._abuse_hit(client_hash, path, flag=_is_abuse_path(path), is_404=True)
self._save()
return client_hash
if _is_abuse_path(path):
self.classify_abuse(ip, client_hash, path, flag=True, is_404=True)
return client_hash
self.not_found_counts[ip] = self.not_found_counts.get(ip, 0) + 1
if self.not_found_counts[ip] >= _ABUSE_404_THRESHOLD:
self.classify_abuse(ip, client_hash, path, flag=True, is_404=True)
return client_hash
return client_hash
def _new_visit( def _new_visit(
self, self,
entry: str, entry: str,
referer: str, referer: str,
key: tuple[str, str], client_hash: bytes,
ip: str = "",
lang: str = "",
country: str = "",
ua: str = "",
utm: dict[str, str] | None = None, utm: dict[str, str] | None = None,
) -> Visit: ) -> Visit:
now = datetime.now(UTC) now = datetime.now(UTC)
@@ -248,50 +559,25 @@ class Store:
start=now, start=now,
entry=entry, entry=entry,
referer=referer, referer=referer,
ip=ip, client=client_hash,
lang=lang,
country=country,
ua=ua,
ua_pretty=_compact_user_agent(ua),
utm=utm or {}, utm=utm or {},
) )
self.data.visits.append(visit) self.data.visits.append(visit)
self.sessions[key] = len(self.data.visits) - 1 self.sessions[client_hash] = len(self.data.visits) - 1
self._count(self.data.site_visits, _bucket(now)) self._count(self.data.site_visits, _bucket(now))
self._count(self.data.views.setdefault(entry, {}), _bucket(now)) self._count(self.data.views.setdefault(entry, {}), _bucket(now))
self._count_transition(referer or "(direct)", entry, now) self._count_transition(referer or "(direct)", entry, now)
return visit return visit
def enrich_visit(
self,
index: int,
*,
host: str = "",
country: str = "",
) -> None:
"""Fill in host/geoip fields on an existing visit after async lookups."""
if index < 0 or index >= len(self.data.visits):
return
visit = self.data.visits[index]
changed = False
if host and not visit.host:
visit.host = host
changed = True
if country:
visit.country = country
changed = True
if changed:
self._save()
def track_entry( def track_entry(
self, self,
referer: str, referer: str,
own_origin: str, own_origin: str,
ip: str, ip: str,
ua: str, ua: str,
entry: str, full_path: str,
query: str = "", accept_language: str = "",
) -> None: ) -> list[bytes]:
"""Stash the entry referer/UTM tags and queue a pending crawler hit. """Stash the entry referer/UTM tags and queue a pending crawler hit.
Nothing is counted here — the client's initial /_a ping starts the Nothing is counted here — the client's initial /_a ping starts the
@@ -302,11 +588,28 @@ class Store:
does not erase an earlier tagged landing. does not erase an earlier tagged landing.
Every document GET is also queued as a pending crawler hit. If a ping Every document GET is also queued as a pending crawler hit. If a ping
from the same (IP, UA) pair arrives within ``_CRAWLER_TIMEOUT``, the from the same client arrives within ``_CRAWLER_TIMEOUT``, the hit is
hit is discarded; otherwise it is flushed to ``data.crawlers``. discarded; otherwise it is flushed to ``data.crawlers``. The
Accept-Language header is stored on the client record immediately;
host/geoip are filled in later by async enrichment.
GETs from IPs already classified as abuse are recorded as abuse hits
with the full request path (query string included).
Returns the client hashes of any hits flushed to persistent storage,
so callers can schedule async enrichment.
""" """
entry = full_path.split("?")[0]
query = full_path.split("?", 1)[1] if "?" in full_path else ""
lang, country = _parse_accept_language(accept_language)
client_hash = self._ensure_client(ip, ua, lang, country=country)
if ip in self.data.abuse_ips:
flushed = self._flush_crawlers()
self._abuse_hit(client_hash, full_path, is_404=False, flag=False)
self._save()
return flushed
now = datetime.now(UTC) now = datetime.now(UTC)
self._flush_crawlers(now) flushed = self._flush_crawlers(now)
if referer: if referer:
origin = _origin(referer) origin = _origin(referer)
if origin is not None and origin != own_origin: if origin is not None and origin != own_origin:
@@ -318,61 +621,98 @@ class Store:
CrawlerHit( CrawlerHit(
start=now, start=now,
entry=entry, entry=entry,
ip=ip, client=client_hash,
ua=ua,
ua_pretty=_compact_user_agent(ua),
referer=self.pending_referers.get(ip, ""), referer=self.pending_referers.get(ip, ""),
query=query, query=query,
) )
) )
return flushed
def _add_read(self, client_hash: bytes, path: str, seconds: int) -> None:
"""Add ``seconds`` of reading time for ``path`` to the current visit."""
if seconds <= 0:
return
index = self.sessions.get(client_hash)
if index is None or index >= len(self.data.visits):
return
visit = self.data.visits[index]
visit.read[path] = visit.read.get(path, 0) + seconds
def ping( def ping(
self, self,
from_: str, from_: str,
to: str, to: str | None,
ip: str, ip: str,
ua: str, ua: str,
accept_language: str = "", accept_language: str = "",
) -> int | None: hide: bool = False,
"""Record a client navigation ping ({from, to} from pagerite.js). read: int = 0,
) -> tuple[int | None, list[bytes]]:
"""Record a client navigation ping ({from, to, read} from pagerite.js).
``to`` is an internal path ("/...") or an https URL for exit links; a
missing/empty ``to`` means the page is being closed and only the
``read`` time should be recorded. The transition is always counted when
``to`` is present; the trail only grows on first sight of a page within
the visit. ``read`` is the active time (seconds) spent on ``from_``.
``to`` is an internal path ("/...") or an https origin for exit
links; anything else is ignored. The transition is always counted;
the trail only grows on first sight of a page within the visit.
A ping with no known session starts a fresh visit, consuming the A ping with no known session starts a fresh visit, consuming the
referer and UTM tags stashed by the document GET if there are any. referer and UTM tags stashed by the document GET if there are any.
Returns the index of the new visit when one is created, so callers ``hide`` is set by admin clients: the ping cancels pending crawler
can enrich it later with non-blocking lookups (host, geoip country). hits as usual, and any existing visit for this client session is
removed from the stats (the admin browsed anonymously before logging
in). Nothing new is recorded.
Pings from IPs classified as abuse are ignored entirely.
Returns the index of the new visit when one is created (or None) and
the client hashes of any crawler hits flushed by this call, so callers
can schedule async enrichment (host, geoip country/city).
""" """
self._flush_crawlers() flushed = self._flush_crawlers()
# A real visitor ping cancels any pending crawler hits from this lang, country = _parse_accept_language(accept_language)
# (IP, UA) pair. client_hash = _client_hash(ip, ua, lang)
if hide:
# Admin ping: cancel pending crawler hits and scrub the session.
self.pending_crawlers = [
hit for hit in self.pending_crawlers if hit.client == client_hash
]
index = self.sessions.pop(client_hash, None)
if index is not None and index < len(self.data.visits):
self._remove_visit(index)
self._save()
return None, flushed
if ip in self.data.abuse_ips:
return None, flushed
# A real visitor ping cancels any pending crawler hits from this client.
self.pending_crawlers = [ self.pending_crawlers = [
hit for hit in self.pending_crawlers if not (hit.ip == ip and hit.ua == ua) hit for hit in self.pending_crawlers if hit.client != client_hash
] ]
fr_path = _internal_path(from_) if from_ else ""
if fr_path and read > 0:
self._add_read(client_hash, fr_path, read)
if not to:
if read > 0:
self._save()
return None, flushed
if to.startswith("/") and not to.startswith("//"): if to.startswith("/") and not to.startswith("//"):
target = _internal_path(to) or "" target = _internal_path(to) or ""
else: else:
target = _origin(to) or "" target = _external_target(to) or ""
if not target or (not to.startswith("/") and target != to): if not target:
return None return None, flushed
key = (ip, ua) index = self.sessions.get(client_hash)
index = self.sessions.get(key) fr = fr_path or "(direct)"
fr = (_internal_path(from_) or "(direct)") if from_ else "(direct)"
if index is None or index >= len(self.data.visits): if index is None or index >= len(self.data.visits):
# No known session: the initial ping of a fresh page load (or # No known session: the initial ping of a fresh page load (or
# missing data after a server restart) — start a visit. # missing data after a server restart) — start a visit.
lang, country = _parse_accept_language(accept_language)
index = len(self.data.visits) index = len(self.data.visits)
self._ensure_client(ip, ua, lang, country=country)
self._new_visit( self._new_visit(
target, target,
self.pending_referers.pop(ip, ""), self.pending_referers.pop(ip, ""),
key, client_hash,
ip=ip,
lang=lang,
country=country,
ua=ua,
utm=self.pending_utms.pop(ip, {}), utm=self.pending_utms.pop(ip, {}),
) )
else: else:
@@ -385,4 +725,5 @@ class Store:
if visit.entry != target and target not in visit.trail: if visit.entry != target and target not in visit.trail:
visit.trail.append(target) visit.trail.append(target)
self._save() self._save()
return index if index is not None and index < len(self.data.visits) else None visit_index = index if index is not None and index < len(self.data.visits) else None
return visit_index, flushed
+143 -26
View File
@@ -57,6 +57,10 @@ ANALYTICS_PATH = Path(
) )
analytics_store = analytics.Store(ANALYTICS_PATH) analytics_store = analytics.Store(ANALYTICS_PATH)
# Live WebSocket clients for the analytics stream.
_analytics_ws_clients: set[WebSocket] = set()
_analytics_broadcast_task: asyncio.Task | None = None
# Repository root from this file's location (pagerite/app.py -> ..). # Repository root from this file's location (pagerite/app.py -> ..).
_REPO_ROOT = Path(__file__).resolve().parent.parent _REPO_ROOT = Path(__file__).resolve().parent.parent
@@ -121,6 +125,26 @@ class GeoIP:
pass pass
return "" return ""
def city(self, ip: str) -> str:
"""City name for ``ip``, or "" when unavailable.
GeoIP sometimes appends district names in parentheses (e.g.
"Berlin (Bezirk Tempelhof-Schöneberg)"); those are stripped before
the value is stored.
"""
if not ip or self._reader is None:
return ""
try:
rec = self._reader.get(ip)
if rec:
city = (rec.get("city") or {}).get("names", {}).get("en", "")
if city:
city = re.sub(r"\s*\([^)]*\)", "", city).strip()
return city
except Exception:
pass
return ""
_geoip = GeoIP() _geoip = GeoIP()
@@ -231,7 +255,9 @@ async def lifespan(_app: FastAPI) -> AsyncIterator[None]:
# Decompress/open the DB-IP MMDB once at startup. Lookups are then # Decompress/open the DB-IP MMDB once at startup. Lookups are then
# read-only and safe to run in background ``to_thread`` workers. # read-only and safe to run in background ``to_thread`` workers.
await asyncio.to_thread(_geoip._load) await asyncio.to_thread(_geoip._load)
analytics_store.subscribe(_schedule_analytics_broadcast)
yield yield
analytics_store.unsubscribe(_schedule_analytics_broadcast)
await kanta.close() await kanta.close()
@@ -587,6 +613,12 @@ def _client_ip(request: Request) -> str:
return forwarded or (request.client.host if request.client else "") return forwarded or (request.client.host if request.client else "")
def _query_suffix(request: Request) -> str:
"""The request's query string as a "?..." suffix, or "" when absent."""
query = str(request.url.query)
return f"?{query}" if query else ""
@lru_cache(maxsize=4096) @lru_cache(maxsize=4096)
def _cached_ptr(ip: str) -> str: def _cached_ptr(ip: str) -> str:
"""Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or "".""" """Reverse-DNS lookup with in-RAM LRU cache. Returns the host name or ""."""
@@ -615,27 +647,75 @@ async def _geoip_country(ip: str) -> str:
return await asyncio.to_thread(_geoip.country, ip) return await asyncio.to_thread(_geoip.country, ip)
async def _enrich_visit(index: int, ip: str) -> None: async def _geoip_city(ip: str) -> str:
"""Run non-blocking reverse-DNS and geoip enrichment for a new visit.""" """Async wrapper around the DB-IP MMDB city lookup."""
if not ip: return await asyncio.to_thread(_geoip.city, ip)
async def _enrich_client(client_hash: bytes) -> None:
"""Run non-blocking reverse-DNS and geoip enrichment for a client."""
client = analytics_store.data.clients.get(client_hash)
if not client or not client.ip:
return return
host = await _lookup_host(ip) host = await _lookup_host(client.ip)
country = await _geoip_country(ip) country = await _geoip_country(client.ip)
analytics_store.enrich_visit(index, host=host, country=country) city = await _geoip_city(client.ip)
analytics_store.enrich_client(client_hash, host=host, country=country, city=city)
def _schedule_client_enrichment(client_hashes: list[bytes]) -> None:
"""Start background host/geoip enrichment for the given client hashes."""
for client_hash in client_hashes:
asyncio.create_task(_enrich_client(client_hash))
async def _broadcast_analytics() -> None:
"""Send the current analytics snapshot to every connected WS client."""
if not _analytics_ws_clients:
return
payload = msgspec.json.encode(analytics_store.data).decode()
closed = set()
for ws in _analytics_ws_clients:
try:
await ws.send_text(payload)
except Exception:
closed.add(ws)
for ws in closed:
_analytics_ws_clients.discard(ws)
async def _debounced_analytics_broadcast() -> None:
"""Wait briefly, then broadcast the latest snapshot once."""
await asyncio.sleep(0.2)
await _broadcast_analytics()
def _schedule_analytics_broadcast() -> None:
"""Schedule a single debounced broadcast, ignoring duplicate triggers."""
global _analytics_broadcast_task
if _analytics_broadcast_task is not None and not _analytics_broadcast_task.done():
return
_analytics_broadcast_task = asyncio.get_running_loop().create_task(
_debounced_analytics_broadcast()
)
class AnalyticsPing(BaseModel): class AnalyticsPing(BaseModel):
"""Navigation ping from pagerite.js (see docs/analytics.md).""" """Navigation ping from pagerite.js (see docs/analytics.md)."""
fr: str = "" fr: str = ""
to: str to: str | None = None
#: 1 from admin clients: scrub the session instead of recording it.
hide: int = 0
#: Active reading time on ``fr`` (ms), if any.
read: int = 0
@app.get("/_a", response_model=None) @app.get("/_a", response_model=None)
async def analytics_page(request: Request) -> HTMLResponse: async def analytics_page(request: Request) -> HTMLResponse:
"""Render the analytics viewer as a normal site page at /_a. """Render the analytics viewer as a normal site page at /_a.
The page itself is public, but the data endpoint (/_api/analytics) stays The page itself is public, but the data stream (/_api/ws/analytics) stays
admin-gated like the rest of /_api, so only authorized users see the admin-gated like the rest of /_api, so only authorized users see the
statistics; others get the viewer with a "could not be loaded" message. statistics; others get the viewer with a "could not be loaded" message.
""" """
@@ -655,42 +735,51 @@ async def analytics_ping(ping: AnalyticsPing, request: Request) -> None:
the response is never delayed by slow DNS or the first MMDB decompress. the response is never delayed by slow DNS or the first MMDB decompress.
""" """
ip = _client_ip(request) ip = _client_ip(request)
index = analytics_store.ping( visit_index, flushed_clients = analytics_store.ping(
ping.fr, ping.fr,
ping.to, ping.to,
ip, ip,
request.headers.get("user-agent", ""), request.headers.get("user-agent", ""),
request.headers.get("accept-language", ""), request.headers.get("accept-language", ""),
hide=bool(ping.hide),
read=ping.read,
) )
if index is not None: if visit_index is not None:
asyncio.create_task(_enrich_visit(index, ip)) visit = analytics_store.data.visits[visit_index]
asyncio.create_task(_enrich_client(visit.client))
_schedule_client_enrichment(flushed_clients)
def _track_entry(path: str, request: Request) -> None: def _track_entry(path: str, request: Request) -> list[bytes]:
"""Stash the referer/UTM tags and queue a pending crawler hit for the GET. """Stash the referer/UTM tags and queue a pending crawler hit for the GET.
Nothing is counted on the GET itself — the client's /_a ping starts the Nothing is counted on the GET itself — the client's /_a ping starts the
visit, so bots and admin browsing never register as visits. visit, so bots never register as visits. (Admin clients ping too, but
with hide=1, which scrubs their session instead of recording it.)
The devserver's health probe (``GET /?from=devserver.py`` from The devserver's health probe (``GET /?from=devserver.py`` from
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be ``127.0.0.1``) is ignored: it is not real traffic and would otherwise be
logged as a crawler hit. The root-path and localhost checks prevent logged as a crawler hit. The root-path and localhost checks prevent
remote visitors from hiding traffic with the same query string. remote visitors from hiding traffic with the same query string.
Returns the client hashes of any pending crawler hits flushed to persistent
storage, so callers can schedule async geoip and reverse-DNS enrichment.
""" """
if ( if (
path == "" path == ""
and str(request.url.query) == "from=devserver.py" and str(request.url.query) == "from=devserver.py"
and _client_ip(request) == "127.0.0.1" and _client_ip(request) == "127.0.0.1"
): ):
return return []
own_origin = f"https://{urlparse(str(request.base_url)).netloc}" own_origin = f"https://{urlparse(str(request.base_url)).netloc}"
analytics_store.track_entry( full_path = f"{request.url.path}{_query_suffix(request)}"
return analytics_store.track_entry(
request.headers.get("referer", ""), request.headers.get("referer", ""),
own_origin, own_origin,
_client_ip(request), _client_ip(request),
request.headers.get("user-agent", ""), request.headers.get("user-agent", ""),
"/" if path == "" else f"/{path}", full_path,
str(request.url.query), request.headers.get("accept-language", ""),
) )
@@ -727,16 +816,23 @@ def _check_reserved(path: str) -> None:
) )
@app.get("/_api/analytics") @app.websocket("/_api/ws/analytics")
async def get_analytics() -> Response: async def analytics_websocket(ws: WebSocket) -> None:
"""The collected visit analytics as JSON (see docs/analytics.md). """Stream the analytics snapshot, then push updates as they happen.
Admin-only via the /_api forward-auth gate, like every management Admin-only via the /_api forward-auth gate, like every management
endpoint. Powers the analytics viewer rendered at /_a. endpoint. Powers the analytics viewer rendered at /_a.
""" """
return Response( await ws.accept()
msgspec.json.encode(analytics_store.data), media_type="application/json" await ws.send_text(msgspec.json.encode(analytics_store.data).decode())
) _analytics_ws_clients.add(ws)
try:
while True:
await ws.receive_text()
except Exception:
pass
finally:
_analytics_ws_clients.discard(ws)
@app.websocket("/_api/ws/editor") @app.websocket("/_api/ws/editor")
@@ -903,9 +999,20 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
placeholder page (nav links point straight at its first child). placeholder page (nav links point straight at its first child).
""" """
path = path.strip("/") path = path.strip("/")
ua = request.headers.get("user-agent", "")
accept_language = request.headers.get("accept-language", "")
if path and _is_reserved(path): if path and _is_reserved(path):
# Invalid slug shape: not a content URL, let FastAPI return its # Invalid slug shape: not a content URL, let FastAPI return its
# built-in 404 instead of rendering an editable article page. # built-in 404 instead of rendering an editable article page.
# Scanner telltales (dotpaths like /.env, *.php) classify the IP
# as abuse in analytics.
client_hash = analytics_store.track_404(
_client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
raise HTTPException(404) raise HTTPException(404)
chain = resolve(data.menu, path) chain = resolve(data.menu, path)
node = chain[-1] if chain else None node = chain[-1] if chain else None
@@ -920,7 +1027,8 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
if request.headers.get("if-none-match") == etag: if request.headers.get("if-none-match") == etag:
return Response(status_code=304) return Response(status_code=304)
if _is_trackable_path(path): if _is_trackable_path(path):
_track_entry(path, request) flushed = _track_entry(path, request)
_schedule_client_enrichment(flushed)
return HTMLResponse( return HTMLResponse(
views.render_page(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html, str(request.base_url).rstrip("/")), views.render_page(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html, str(request.base_url).rstrip("/")),
headers={ headers={
@@ -933,7 +1041,8 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
# Category label without a landing page: placeholder with the pen # Category label without a landing page: placeholder with the pen
# to create it (404 — no page here, but the node is real). # to create it (404 — no page here, but the node is real).
if _is_trackable_path(path): if _is_trackable_path(path):
_track_entry(path, request) flushed = _track_entry(path, request)
_schedule_client_enrichment(flushed)
return HTMLResponse( return HTMLResponse(
views.render_category(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html), views.render_category(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html),
404, 404,
@@ -949,5 +1058,13 @@ async def show_page(request: Request, path: str) -> HTMLResponse | Response:
if item.published: if item.published:
return RedirectResponse(f"/{slug}") return RedirectResponse(f"/{slug}")
if _is_trackable_path(path): if _is_trackable_path(path):
_track_entry(path, request) client_hash = analytics_store.track_404(
_client_ip(request),
ua,
f"/{path}{_query_suffix(request)}",
accept_language,
)
asyncio.create_task(_enrich_client(client_hash))
flushed = _track_entry(path, request)
_schedule_client_enrichment(flushed)
return HTMLResponse(views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html), 404) return HTMLResponse(views.render_not_found(data.menu, path, data.brand, data.custom_css, data.theme, data.favicon, data.brand_html), 404)
+5 -3
View File
@@ -29,6 +29,7 @@ Where to go next:
- The [docs](/docs/editing) section explains how to edit this site and shows every supported Markdown feature, source and result side by side. - The [docs](/docs/editing) section explains how to edit this site and shows every supported Markdown feature, source and result side by side.
- The [showcase](/showcase/gallery) section shows what finished pages can look like: image positioning, banners, a long read. - The [showcase](/showcase/gallery) section shows what finished pages can look like: image positioning, banners, a long read.
- Click the 🖊️ pen on any page to open the editor, and the ⚙️ pen for site settings and the structure tree. - Click the 🖊️ pen on any page to open the editor, and the ⚙️ pen for site settings and the structure tree.
- Elsewhere on the web: [![xkcd 927: Standards](https://imgs.xkcd.com/comics/standards.png "xkcd 927: Standards"){width=240}](https://xkcd.com/927/) — a cautionary tale about adding one more standard.
![Abstract waves](waves.svg "Generated SVG artwork, attached to this page"){width=420} ![Abstract waves](waves.svg "Generated SVG artwork, attached to this page"){width=420}
@@ -68,15 +69,15 @@ Every feature below is shown twice: first the Markdown source, then how it rende
### A subsection ### A subsection
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a *Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a
[link to the front page](/). Plain URLs become links automatically: [link to the front page](/). An image that links to its page:
https://example.com — and a hard line break [![xkcd 1179: ISO 8601](https://imgs.xkcd.com/comics/iso_8601.png "xkcd 1179: ISO 8601"){width=240}](https://xkcd.com/1179/) — and a hard line break
is just a newline. is just a newline.
``` ```
## A section heading ## A section heading
### A subsection ### A subsection
*Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a [link to the front page](/). Plain URLs become links automatically: https://example.com — and a hard line break *Emphasis*, **strong**, ~~strikethrough~~, `inline code`, and a [link to the front page](/). An image that links to its page: [![xkcd 1179: ISO 8601](https://imgs.xkcd.com/comics/iso_8601.png "xkcd 1179: ISO 8601"){width=240}](https://xkcd.com/1179/) — and a hard line break
is just a newline. is just a newline.
## Lists and quotes ## Lists and quotes
@@ -354,6 +355,7 @@ This site runs on **Pagerite**: FastAPI + html5tagger + kanta, with content writ
- [How to edit this site](/docs/editing) - [How to edit this site](/docs/editing)
- [Markdown features](/docs/markdown/basics) - [Markdown features](/docs/markdown/basics)
- [The showcase](/showcase/gallery) - [The showcase](/showcase/gallery)
- [![xkcd 2347: Dependency](https://imgs.xkcd.com/comics/dependency.png "xkcd 2347: Dependency"){width=240}](https://xkcd.com/2347/) — a small comic about small dependencies
*Replace this page with whatever your site is about.* *Replace this page with whatever your site is about.*
""" """
+2 -3
View File
@@ -20,6 +20,7 @@ dependencies = [
"fastapi-vue>=1.3.1", "fastapi-vue>=1.3.1",
"fastapi[standard]>=0.141.1", "fastapi[standard]>=0.141.1",
"html5tagger>=2.0.0", "html5tagger>=2.0.0",
"httpx>=0.28.1",
"kanta>=0.8.1", "kanta>=0.8.1",
"markdown-it-py>=4.2.0", "markdown-it-py>=4.2.0",
"maxminddb>=3.1.1", "maxminddb>=3.1.1",
@@ -35,9 +36,7 @@ pagerite = "pagerite.__main__:main"
Repository = "https://git.zi.fi/LeoVasanko/pagerite" Repository = "https://git.zi.fi/LeoVasanko/pagerite"
[dependency-groups] [dependency-groups]
dev = [ dev = []
"httpx>=0.28.1",
]
[tool.hatch.version] [tool.hatch.version]
source = "vcs" source = "vcs"
+2 -2
View File
@@ -19,8 +19,8 @@ from devutil import (
setup_vite, setup_vite,
) )
DEFAULT_VITE_PORT = 3100 DEFAULT_VITE_PORT = 8200
DEFAULT_DEV_PORT = 3200 DEFAULT_DEV_PORT = 8210
HEALTH = "/?from=devserver.py" HEALTH = "/?from=devserver.py"
+664 -143
View File
@@ -6,20 +6,18 @@
# "playwright>=1.45.0", # "playwright>=1.45.0",
# ] # ]
# /// # ///
"""Generate fake browser visits and crawler hits for a Pagerite site. """Generate fake browser visits, crawler hits, and abuse scans for a Pagerite site.
The script drives a real Chromium browser with Playwright, clicking visible Browser sessions (ordinary users) come from realistic residential IPv4 and IPv6
internal links so the site's own analytics JavaScript records normal visits addresses and stay mostly stable; an IPv6 host part may rotate once mid-session,
(POST /_a). Browser sessions and crawler GETs send a small rotating pool of and an IPv4 session may switch to another residential address. Crawler hits come
real public IPs in X-Forwarded-For, so the backend can reverse-DNS and GeoIP from datacenter IPs, with each crawler profile paired to a matching provider IP
them instead of seeing every hit as 127.0.0.1. when possible. Abuse scanners fire bursts of vulnerability probes from pinned
datacenter IPs.
Sessions start with a Poisson inter-arrival delay (``--arrival-rate``) to
spread traffic out a little, while still keeping the overall run fast.
Run against a local dev server, e.g.: Run against a local dev server, e.g.:
uv run scripts/fake_traffic.py http://localhost:3200 -b 8 -c 20 uv run scripts/fake_traffic.py http://localhost:3200
Repeat whenever you want more traffic; each run appends new events to the Repeat whenever you want more traffic; each run appends new events to the
site's analytics file. site's analytics file.
@@ -36,7 +34,7 @@ from collections.abc import Sequence
from dataclasses import dataclass from dataclasses import dataclass
from datetime import UTC, datetime from datetime import UTC, datetime
from typing import Any from typing import Any
from urllib.parse import urljoin, urlparse from urllib.parse import urlencode, urljoin, urlparse
import httpx import httpx
@@ -56,6 +54,7 @@ class BrowserProfile:
class CrawlerProfile: class CrawlerProfile:
name: str name: str
user_agent: str user_agent: str
ip: str
BROWSER_PROFILES: list[BrowserProfile] = [ BROWSER_PROFILES: list[BrowserProfile] = [
@@ -92,37 +91,419 @@ CRAWLER_PROFILES: list[CrawlerProfile] = [
CrawlerProfile( CrawlerProfile(
"googlebot", "googlebot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36", "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
"66.249.64.66", # US, Google
), ),
CrawlerProfile( CrawlerProfile(
"bingbot", "bingbot",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36", "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
"40.77.167.0", # US, Microsoft
), ),
CrawlerProfile( CrawlerProfile(
"duckduckbot", "DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)" "duckduckbot",
"DuckDuckBot/1.1; (+http://duckduckgo.com/duckduckbot.html)",
"95.217.0.1", # Germany, Hetzner VPS
),
CrawlerProfile(
"curl",
"curl/8.5.0",
"139.162.0.1", # Singapore, Linode VPS
), ),
CrawlerProfile("curl", "curl/8.5.0"),
] ]
# Small pool of real public resolver IPs. They have real reverse-DNS and GeoIP # Residential IPv4 addresses and IPv6 /64 prefixes used for ordinary browser
# entries, and cycling through a handful avoids hammering DNS during traffic # sessions. IPv6 entries keep the network part stable and randomise only the
# generation. # host part; the host may rotate once mid-session.
SOURCE_IPS: list[str] = [ RESIDENTIAL_SOURCE_IPS: list[str] = [
"8.8.8.8", # Residential IPv4
"1.1.1.1", "91.154.140.209", # Finland, Elisa
"9.9.9.9", "84.143.145.207", # Germany, Deutsche Telekom
"208.67.222.222", "220.165.255.254", # China, Chinanet / China Telecom
"185.228.168.9", "84.235.83.162", # Saudi Arabia, SaudiNet / STC
"94.140.14.14", # Residential IPv6 /64 prefixes
"2a02:8109:ac82:6f0c::/64", # Germany, Deutsche Telekom
"240e:45d:1e60:5b0::/64", # China, China Telecom
"2409:8904:6720:4123::/64", # China, China Unicom
] ]
# Concrete datacenter IPs used for abuse scanner bursts. They stay pinned for
# the whole scan burst.
# Index 0 randomises its UA per request, index 1 uses a fixed browser UA,
# and index 2 uses a fixed crawler UA.
ABUSE_SOURCE_IPS: list[str] = [
"45.63.0.12", # US, Vultr VPS
"138.197.0.89", # US, DigitalOcean / Cloudways
"2a01:4f8:0:2::1234", # Germany, Hetzner VPS
]
# Paths commonly probed by attackers looking for exposed config, admin panels,
# version control, credentials, backups, or debug endpoints.
SUSPICIOUS_PATHS: list[str] = [
"/.env",
"/env",
"/.env.local",
"/env.development",
"/config",
"/config.json",
"/config.yaml",
"/config.yml",
"/configuration.json",
"/configuration.yaml",
"/configuration.yml",
"/settings.json",
"/settings.yaml",
"/settings.yml",
"/app.config",
"/appsettings.json",
"/appsettings.Development.json",
"/credentials",
"/credentials.json",
"/secrets",
"/secrets.json",
"/.aws/credentials",
"/.ssh/id_rsa",
"/id_rsa",
"/id_rsa.pub",
"/known_hosts",
"/sftp-config.json",
"/admin",
"/administrator",
"/adminer.php",
"/login",
"/signin",
"/auth/login",
"/api/login",
"/api/.env",
"/api/config",
"/api/v1/config",
"/api/v2/config",
"/webhook",
"/webhooks",
"/callback",
"/proxy",
"/image",
"/images",
"/preview",
"/download",
"/downloads",
"/log",
"/logs",
"/debug",
"/trace",
"/phpinfo.php",
"/info.php",
"/phpmyadmin",
"/pma",
"/myadmin",
"/phpMyAdmin",
"/wp-admin",
"/wp-login.php",
"/wp-config.php",
"/xmlrpc.php",
"/wp-json/wp/v2/users",
"/.git/config",
"/.git/HEAD",
"/git/config",
"/swagger-ui.html",
"/v2/api-docs",
"/actuator/env",
"/actuator/health",
"/actuator/configprops",
"/server-status",
"/.htaccess",
"/web.config",
"/package.json",
"/composer.json",
"/vendor/autoload.php",
"/docker-compose.yml",
"/Dockerfile",
"/manage",
"/console",
"/manager",
"/manager/html",
"/metrics",
"/prometheus",
"/healthz",
"/_api",
"/api",
"/api/v1/",
"/api/v2/",
"/graphql",
"/query",
"/feed",
"/rss",
"/_debug",
"/test",
"/testing",
"/tmp",
"/temp",
"/backup",
"/backups",
"/dump",
"/dumps",
"/sql",
"/db",
"/database",
"/dump.sql",
"/backup.sql",
"/db.sql",
"/backup.zip",
"/backup.tar.gz",
"/site.zip",
"/site.tar.gz",
"/source.zip",
"/src.zip",
"/upload",
"/uploads",
"/import",
"/export",
"/token",
"/tokens",
"/oauth",
"/oauth2",
"/openid",
"/jwks",
"/keys",
"/key",
"/private",
"/public",
]
# Realistic external referers. Most sessions arrive with a generic referer;
# a subset carries matching UTM tags on the landing URL.
PLAIN_REFERRERS: list[str] = [
"https://example.com/",
"https://somedomain.com/",
"https://another-site.org/",
"https://friend-site.net/",
]
# (referer origin, utm parameter dict) pairs used for tagged traffic.
TAGGED_REFERRERS: list[tuple[str, dict[str, str]]] = [
("https://chatgpt.com/", {"utm_source": "chatgpt.com"}),
("https://www.google.com/", {"utm_source": "google", "utm_medium": "organic"}),
("https://twitter.com/", {"utm_source": "twitter", "utm_medium": "social"}),
("https://www.linkedin.com/", {"utm_source": "linkedin", "utm_medium": "social"}),
("https://github.com/", {"utm_source": "github", "utm_medium": "referral"}),
("https://news.ycombinator.com/", {"utm_source": "hackernews", "utm_medium": "referral"}),
("https://www.reddit.com/", {"utm_source": "reddit", "utm_medium": "social"}),
("https://medium.com/", {"utm_source": "medium", "utm_medium": "referral"}),
("https://www.producthunt.com/", {"utm_source": "producthunt", "utm_medium": "referral"}),
]
# Fraction of referered sessions that also carry UTM tags.
UTM_RATE = 0.25
# Innocent-looking paths that do not exist on a Pagerite site. Hitting many of
# these from a single IP is itself a telltale of a spray-and-pray scanner.
NORMAL_404_PATHS: list[str] = [
"/about",
"/about-us",
"/services",
"/products",
"/contact",
"/contact-us",
"/team",
"/careers",
"/jobs",
"/pricing",
"/features",
"/demo",
"/trial",
"/docs",
"/documentation",
"/api-docs",
"/support",
"/help",
"/faq",
"/knowledge-base",
"/terms",
"/terms-of-service",
"/privacy",
"/privacy-policy",
"/legal",
"/blog",
"/news",
"/articles",
"/press",
"/events",
"/webinars",
"/podcast",
"/videos",
"/resources",
"/whitepapers",
"/case-studies",
"/customers",
"/clients",
"/testimonials",
"/reviews",
"/partners",
"/integrations",
"/api-reference",
"/developers",
"/status",
"/security",
"/trust",
"/compliance",
"/gdpr",
"/ccpa",
"/sitemap",
"/archive",
"/tags",
"/categories",
"/search",
"/users",
"/accounts",
"/dashboard",
"/profile",
"/settings",
"/preferences",
"/notifications",
"/messages",
"/inbox",
"/calendar",
"/reports",
"/analytics",
"/billing",
"/invoice",
"/orders",
"/cart",
"/checkout",
"/store",
"/shop",
"/home",
"/main",
"/start",
"/welcome",
"/intro",
"/overview",
"/summary",
"/portfolio",
"/projects",
"/work",
"/solutions",
]
ABUSE_USER_AGENTS: list[str] = [
# Desktop browsers
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Safari/605.1.15",
"Mozilla/5.0 (X11; Linux x86_64; rv:130.0) Gecko/20100101 Firefox/130.0",
"Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:130.0) Gecko/20100101 Firefox/130.0",
"Mozilla/5.0 (Linux; Android 14; SM-S918B) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36",
"Mozilla/5.0 (iPhone; CPU iPhone OS 17_5 like Mac OS X) AppleWebKit/605.1.15 "
"(KHTML, like Gecko) Version/17.5 Mobile/15E148 Safari/604.1",
# Well-known crawlers / bots
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; "
"+http://www.google.com/bot.html) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; "
"+http://www.bing.com/bingbot.htm) Chrome/128.0.0.0 Safari/537.36",
"Mozilla/5.0 (compatible; DuckDuckBot/1.1; +http://duckduckgo.com/duckduckbot.html)",
"Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)",
"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0.0.0 Mobile Safari/537.36 "
"(compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
"Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)",
"Mozilla/5.0 (compatible; DotBot/1.2; +https://opensiteexplorer.org/dotbot; help@moz.com)",
"Mozilla/5.0 (compatible; SemrushBot/7~bl; +http://www.semrush.com/bot.html)",
"Mozilla/5.0 (compatible; AhrefsBot/7.0; +http://ahrefs.com/robot/)",
# Social / service fetchers
"facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)",
"Twitterbot/1.0",
"LinkedInBot/1.0 (compatible; Mozilla/5.0; Apache-HttpClient +http://www.linkedin.com)",
"Slackbot-LinkExpanding 1.0 (+https://api.slack.com/robots)",
"WhatsApp/2.23.20.0",
# Command-line / library clients
"curl/8.5.0",
"Wget/1.21.4 (linux-gnu)",
"python-requests/2.32.3",
"Go-http-client/1.1",
"Node.js/20.5.1",
]
def _random_ipv6_host(prefix: str) -> str:
"""Return a concrete address within an IPv6 /64 prefix.
The host part is generated randomly, mimicking a fresh OS privacy address.
The input prefix must end in ``::/64`` (e.g. ``2a02:8109:ac82:6f0c::/64``).
"""
if "/" not in prefix:
return prefix
base, mask = prefix.split("/")
if mask != "64":
raise ValueError(f"only /64 IPv6 prefixes are supported, got {prefix!r}")
if base.endswith("::"):
base = base[:-2]
host = ":".join(f"{random.randint(0, 0xffff):04x}" for _ in range(4))
return f"{base}:{host}"
def _concretize_ip(entry: str) -> str:
"""Return a concrete IP address; randomise the host part for IPv6 /64 prefixes."""
if ":" in entry and "/" in entry:
return _random_ipv6_host(entry)
return entry
class _SessionIP:
"""Stable IP for a browser session, with one optional mid-session rotation.
IPv6 prefixes get a fresh random host part; IPv4 addresses are swapped for
another address from the residential pool.
"""
def __init__(self, entry: str, pool: Sequence[str]):
self.entry = entry
self.pool = pool
self._value = _concretize_ip(entry)
def current(self) -> str:
return self._value
def rotate(self) -> None:
if ":" in self.entry and "/" in self.entry:
self._value = _random_ipv6_host(self.entry)
return
# IPv4: switch to another IPv4 address from the residential pool.
for _ in range(20):
candidate_entry = random.choice(self.pool)
if ":" in candidate_entry and "/" in candidate_entry:
continue
candidate = _concretize_ip(candidate_entry)
if candidate != self._value:
self._value = candidate
return
def _sleep(base: float, jitter: float) -> None: def _sleep(base: float, jitter: float) -> None:
time.sleep(max(0.0, base + random.uniform(-jitter, jitter))) time.sleep(max(0.0, base + random.uniform(-jitter, jitter)))
def _source_ip(index: int) -> str: def _normalize_url(url: str) -> str:
"""Pick one of the small pool of real public IPs.""" """Return a usable base URL, adding missing scheme/host/port parts.
return SOURCE_IPS[index % len(SOURCE_IPS)]
- bare ``:PORT`` becomes ``http://localhost:PORT``
- missing scheme becomes ``http://``
- otherwise returned as-is
Raises ``ValueError`` when the result is not a valid http(s) URL.
"""
raw = url.strip()
if not raw:
raise ValueError("empty URL")
if raw.startswith(":"):
raw = f"http://localhost{raw}"
elif raw.isdigit():
raw = f"http://localhost:{raw}"
elif not raw.startswith(("http://", "https://")):
raw = f"http://{raw}"
parsed = urlparse(raw)
if parsed.scheme not in ("http", "https") or not parsed.netloc:
raise ValueError(f"invalid URL: {url!r}")
return raw
def _poisson_wait(rate: float) -> float: def _poisson_wait(rate: float) -> float:
@@ -132,31 +513,39 @@ def _poisson_wait(rate: float) -> float:
return random.expovariate(rate) return random.expovariate(rate)
def _collect_links(page: Any) -> list[dict[str, Any]]: def _collect_links(page: Any, include_external: bool = False) -> list[dict[str, Any]]:
"""Return internal links from the current page, excluding the current page.""" """Return links from the current page, excluding the current page.
Internal links stay on the site; external links are real https URLs found
in the page content and are marked with ``external: true``.
"""
return page.evaluate( return page.evaluate(
"""() => { """(includeExternal) => {
const loc = new URL(location.href); const loc = new URL(location.href);
return Array.from(document.querySelectorAll('a[href]')) const out = [];
.filter(a => { for (const a of document.querySelectorAll('a[href]')) {
try { try {
const u = new URL(a.href); const u = new URL(a.href);
return u.origin === loc.origin
&& !u.pathname.startsWith('/_')
&& !u.pathname.startsWith('/auth')
&& u.pathname !== '/favicon.ico'
&& u.pathname !== loc.pathname;
} catch { return false; }
})
.map(a => {
const rect = a.getBoundingClientRect(); const rect = a.getBoundingClientRect();
return { const item = {
href: a.href, href: a.href,
text: (a.innerText || a.title || '').trim().slice(0, 60), text: (a.innerText || a.title || '').trim().slice(0, 60),
visible: !!(rect.width && rect.height && rect.top < window.innerHeight && rect.bottom > 0), visible: !!(rect.width && rect.height && rect.top < window.innerHeight && rect.bottom > 0),
}; };
}); if (u.origin === loc.origin
}""" && !u.pathname.startsWith('/_')
&& !u.pathname.startsWith('/auth')
&& u.pathname !== '/favicon.ico'
&& u.pathname !== loc.pathname) {
out.push(item);
} else if (includeExternal && u.protocol === 'https:' && u.origin !== loc.origin) {
out.push({ ...item, external: true });
}
} catch { /* ignore malformed hrefs */ }
}
return out;
}""",
include_external,
) )
@@ -197,37 +586,75 @@ def _run_browser_session(
paths: Sequence[str], paths: Sequence[str],
profile: BrowserProfile, profile: BrowserProfile,
session_index: int, session_index: int,
max_clicks: int, ip_entry: str,
stay: tuple[float, float],
headless: bool,
fake_ip: str,
) -> dict[str, Any]: ) -> dict[str, Any]:
from playwright.sync_api import sync_playwright from playwright.sync_api import sync_playwright
MAX_CLICKS = 6
STAY = (2.0, 6.0)
HEADLESS = True
REFERER_RATE = 0.75
INCLUDE_EXTERNAL = True
ip_provider = _SessionIP(ip_entry, RESIDENTIAL_SOURCE_IPS)
ips_used: list[str] = [ip_provider.current()]
trail: list[str] = [] trail: list[str] = []
start_time = datetime.now(UTC) start_time = datetime.now(UTC)
try: try:
with sync_playwright() as p: with sync_playwright() as p:
browser = p.chromium.launch( browser = p.chromium.launch(
headless=headless, headless=HEADLESS,
args=["--no-sandbox", "--disable-dev-shm-usage"], args=["--no-sandbox", "--disable-dev-shm-usage"],
) )
extra_headers = {
"X-Forwarded-For": ip_provider.current(),
"Accept-Language": profile.accept_language,
}
# Most sessions arrive from an external origin; some are direct.
# A subset of referered sessions carries realistic UTM tags on the
# landing URL; the referer origin is paired with the UTM source.
tagged: dict[str, str] = {}
if random.random() < REFERER_RATE:
if random.random() < UTM_RATE:
referer, tagged = random.choice(TAGGED_REFERRERS)
else:
referer = random.choice(PLAIN_REFERRERS)
extra_headers["Referer"] = referer
context = browser.new_context( context = browser.new_context(
user_agent=profile.user_agent, user_agent=profile.user_agent,
viewport={"width": profile.viewport[0], "height": profile.viewport[1]}, viewport={"width": profile.viewport[0], "height": profile.viewport[1]},
extra_http_headers={ extra_http_headers=extra_headers,
"X-Forwarded-For": fake_ip,
"Accept-Language": profile.accept_language,
},
) )
page = context.new_page() page = context.new_page()
# Update X-Forwarded-For per request; the value stays stable unless we
# explicitly rotate it once mid-session.
def _route_handler(route, request):
headers = dict(request.headers)
headers["X-Forwarded-For"] = ip_provider.current()
ips_used.append(headers["X-Forwarded-For"])
route.continue_(headers=headers)
page.route("**/*", _route_handler)
# Pick one point during the session to emulate an IP rotation.
rotate_at = random.randint(0, MAX_CLICKS - 1) if MAX_CLICKS > 0 else -1
entry = random.choice(paths) if paths else "/" entry = random.choice(paths) if paths else "/"
page.goto(urljoin(base, entry), wait_until="networkidle") landing = urljoin(base, entry)
if tagged:
sep = "&" if "?" in landing else "?"
landing += sep + urlencode(tagged)
page.goto(landing, wait_until="networkidle")
trail.append(page.url) trail.append(page.url)
for _ in range(max_clicks): for click_idx in range(MAX_CLICKS):
_sleep(random.uniform(*stay) / 2, 0.3) _sleep(random.uniform(*STAY) / 2, 0.3)
links = _collect_links(page) if click_idx == rotate_at:
ip_provider.rotate()
ips_used.append(ip_provider.current())
logger.debug("rotated session IP to %s", ip_provider.current())
links = _collect_links(page, INCLUDE_EXTERNAL)
visible = [item for item in links if item.get("visible")] visible = [item for item in links if item.get("visible")]
if not visible: if not visible:
visible = links visible = links
@@ -242,16 +669,23 @@ def _run_browser_session(
ok = _click_link(page, alt) ok = _click_link(page, alt)
if not ok: if not ok:
break break
if link.get("external"):
# Outbound navigation: the analytics exit ping is already
# in flight. Record the external URL and end the session.
trail.append(page.url)
_sleep(0.5, 0.2)
break
page.wait_for_load_state("networkidle") page.wait_for_load_state("networkidle")
trail.append(page.url) trail.append(page.url)
_sleep(random.uniform(*stay), 0.5) _sleep(random.uniform(*STAY), 0.5)
browser.close() browser.close()
return { return {
"profile": profile.name, "profile": profile.name,
"entry": entry, "entry": entry,
"ip": fake_ip, "ip": ips_used[0],
"ips_seen": len(set(ips_used)),
"pages": len(trail), "pages": len(trail),
"trail": [urlparse(u).path or "/" for u in trail], "trail": [urlparse(u).path or "/" for u in trail],
"duration": (datetime.now(UTC) - start_time).total_seconds(), "duration": (datetime.now(UTC) - start_time).total_seconds(),
@@ -265,12 +699,10 @@ def _run_crawler_hit(
base: str, base: str,
paths: Sequence[str], paths: Sequence[str],
profile: CrawlerProfile, profile: CrawlerProfile,
profile_index: int,
session_index: int,
) -> dict[str, Any]: ) -> dict[str, Any]:
path = random.choice(paths) if paths else "/" path = random.choice(paths) if paths else "/"
url = urljoin(base, path) url = urljoin(base, path)
fake_ip = _source_ip(session_index) fake_ip = profile.ip
headers = { headers = {
"User-Agent": profile.user_agent, "User-Agent": profile.user_agent,
"X-Forwarded-For": fake_ip, "X-Forwarded-For": fake_ip,
@@ -290,49 +722,99 @@ def _run_crawler_hit(
return {"profile": profile.name, "path": path, "error": str(exc)} return {"profile": profile.name, "path": path, "error": str(exc)}
def _abuse_ua() -> str:
"""Return a randomized, syntactically valid user agent for an abuse scan."""
return random.choice(ABUSE_USER_AGENTS)
def _run_abuse_scanner(base: str, ip_index: int) -> dict[str, Any]:
"""Fire a burst of vulnerability probes from a single fake IP.
Scanner 0 randomises its user agent every request, scanner 1 uses a fixed
browser UA, and scanner 2 uses a fixed crawler UA.
"""
ip_entry = ABUSE_SOURCE_IPS[ip_index % len(ABUSE_SOURCE_IPS)]
if ":" in ip_entry and "/" in ip_entry:
fake_ip = _random_ipv6_host(ip_entry)
else:
fake_ip = ip_entry
MIN_HITS = 15
MAX_HITS = 25
total_hits = random.randint(MIN_HITS, MAX_HITS)
# Ensure the burst contains both telltales: suspicious paths and more
# than ten normal-looking 404 paths.
suspicious_count = max(5, total_hits // 3)
normal_count = total_hits - suspicious_count
if normal_count < 11:
normal_count = 11
suspicious_count = max(3, total_hits - normal_count)
paths = random.choices(SUSPICIOUS_PATHS, k=suspicious_count) + random.choices(
NORMAL_404_PATHS, k=normal_count
)
random.shuffle(paths)
ua_mode = ip_index % 3
if ua_mode == 0:
get_ua = _abuse_ua
elif ua_mode == 1:
def get_ua() -> str:
return BROWSER_PROFILES[0].user_agent
else:
def get_ua() -> str:
return CRAWLER_PROFILES[0].user_agent
scan_results: list[dict[str, Any]] = []
with httpx.Client(follow_redirects=True, timeout=15.0) as client:
for path in paths:
headers = {
"User-Agent": get_ua(),
"X-Forwarded-For": fake_ip,
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": random.choice(
["en-US,en;q=0.9", "en-GB,en;q=0.8", "en;q=0.7"]
),
}
try:
r = client.get(urljoin(base, path), headers=headers)
scan_results.append(
{"path": path, "status": r.status_code, "ua": headers["User-Agent"]}
)
except Exception as exc: # noqa: BLE001
scan_results.append({"path": path, "error": str(exc)})
_sleep(0.15, 0.1)
return {
"scanner": ip_index + 1,
"ip": fake_ip,
"hits": len(scan_results),
"results": scan_results,
}
def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace: def _parse_args(argv: Sequence[str] | None) -> argparse.Namespace:
parser = argparse.ArgumentParser( parser = argparse.ArgumentParser(
description="Generate fake traffic for a Pagerite site.", description="Generate fake traffic for a Pagerite site.",
formatter_class=argparse.ArgumentDefaultsHelpFormatter, formatter_class=argparse.ArgumentDefaultsHelpFormatter,
) )
parser.add_argument("url", help="Base URL of the Pagerite site")
parser.add_argument( parser.add_argument(
"-b", "url",
"--browsers", nargs="?",
type=int, default="http://localhost:8200",
default=5, help="Base URL of the Pagerite site (default: http://localhost:8200). "
help="Number of simulated browser sessions", "A bare :PORT or PORT is treated as http://localhost:PORT; a "
"missing scheme defaults to http://.",
) )
parser.add_argument( parser.add_argument(
"-c", "--crawlers", type=int, default=10, help="Number of crawler HTTP GETs" "-t",
) "--duration",
parser.add_argument(
"--max-clicks",
type=int,
default=6,
help="Max internal link clicks per browser session",
)
parser.add_argument(
"--stay",
type=float, type=float,
nargs=2, default=60.0,
default=[2.0, 6.0], metavar="SECONDS",
metavar=("MIN", "MAX"), help="Rough maximum time to generate traffic (0 runs one preset batch)",
help="Seconds to stay on a page before clicking again",
) )
parser.add_argument(
"--headless",
action=argparse.BooleanOptionalAction,
default=True,
help="Run browsers headlessly",
)
parser.add_argument(
"--arrival-rate",
type=float,
default=1.0,
help="Average arrivals per second (Poisson). 0 disables inter-arrival waits",
)
parser.add_argument("--seed", type=int, default=None, help="Random seed")
parser.add_argument("-v", "--verbose", action="store_true", help="Debug logging") parser.add_argument("-v", "--verbose", action="store_true", help="Debug logging")
return parser.parse_args(argv) return parser.parse_args(argv)
@@ -342,8 +824,11 @@ def main(argv: Sequence[str] | None = None) -> int:
if args.verbose: if args.verbose:
logger.setLevel(logging.DEBUG) logger.setLevel(logging.DEBUG)
random.seed(args.seed) try:
base = args.url.rstrip("/") base = _normalize_url(args.url).rstrip("/")
except ValueError as exc:
logger.error("%s", exc)
return 2
# Discover content paths from the public page tree if we can. # Discover content paths from the public page tree if we can.
paths: list[str] = [] paths: list[str] = []
@@ -357,59 +842,95 @@ def main(argv: Sequence[str] | None = None) -> int:
paths = ["/"] paths = ["/"]
logger.info( logger.info(
"Generating fake traffic against %s (%d content paths, %d browsers, %d crawlers)", "Generating fake traffic against %s (%d content paths, duration=%ss)",
base, base,
len(paths), len(paths),
args.browsers, args.duration,
args.crawlers,
) )
results: list[dict[str, Any]] = [] results: list[dict[str, Any]] = []
arrival_rate = 1.0
for i in range(args.browsers): def _wait() -> None:
if i > 0: wait = _poisson_wait(arrival_rate)
wait = _poisson_wait(args.arrival_rate) logger.debug("waiting %.2fs before next session", wait)
logger.debug("waiting %.2fs before next browser session", wait) time.sleep(wait)
time.sleep(wait)
profile = random.choice(BROWSER_PROFILES)
fake_ip = _source_ip(i)
logger.info(
"[%d/%d] browser session: %s (ip=%s)",
i + 1,
args.browsers,
profile.name,
fake_ip,
)
result = _run_browser_session(
base,
paths,
profile,
i,
args.max_clicks,
(args.stay[0], args.stay[1]),
args.headless,
fake_ip,
)
results.append(result)
logger.debug(" trail: %s", result.get("trail", []))
for i in range(args.crawlers): if args.duration <= 0:
if i > 0: # One preset batch.
wait = _poisson_wait(args.arrival_rate) for i in range(5):
logger.debug("waiting %.2fs before next crawler hit", wait) if i > 0:
time.sleep(wait) _wait()
profile_index = i % len(CRAWLER_PROFILES) profile = random.choice(BROWSER_PROFILES)
profile = CRAWLER_PROFILES[profile_index] ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
fake_ip = _source_ip(i) logger.info(
logger.info( "browser session: %s (ip=%s)",
"[%d/%d] crawler hit: %s (ip=%s)", profile.name,
i + 1, _concretize_ip(ip_entry),
args.crawlers, )
profile.name, result = _run_browser_session(base, paths, profile, i, ip_entry)
fake_ip, results.append(result)
) logger.debug(" trail: %s", result.get("trail", []))
result = _run_crawler_hit(base, paths, profile, profile_index, i)
results.append(result) for i in range(10):
if i > 0:
_wait()
profile = random.choice(CRAWLER_PROFILES)
logger.info(
"crawler hit: %s (ip=%s)",
profile.name,
profile.ip,
)
result = _run_crawler_hit(base, paths, profile)
results.append(result)
for i in range(3):
if i > 0:
_wait()
ip_entry = ABUSE_SOURCE_IPS[i % len(ABUSE_SOURCE_IPS)]
logger.info("abuse scanner: %s", ip_entry)
result = _run_abuse_scanner(base, i)
results.append(result)
logger.debug(
" hits: %s", [r.get("path") for r in result.get("results", [])]
)
else:
deadline = time.time() + args.duration
session_index = 0
while time.time() < deadline:
if session_index > 0:
_wait()
phase = session_index % 3
if phase == 0:
profile = random.choice(BROWSER_PROFILES)
ip_entry = random.choice(RESIDENTIAL_SOURCE_IPS)
logger.info(
"browser session: %s (ip=%s)",
profile.name,
_concretize_ip(ip_entry),
)
result = _run_browser_session(
base, paths, profile, session_index, ip_entry
)
logger.debug(" trail: %s", result.get("trail", []))
elif phase == 1:
profile = random.choice(CRAWLER_PROFILES)
logger.info(
"crawler hit: %s (ip=%s)",
profile.name,
profile.ip,
)
result = _run_crawler_hit(base, paths, profile)
else:
ip_entry = ABUSE_SOURCE_IPS[session_index % len(ABUSE_SOURCE_IPS)]
logger.info("abuse scanner: %s", ip_entry)
result = _run_abuse_scanner(base, session_index // 3)
logger.debug(
" hits: %s",
[r.get("path") for r in result.get("results", [])],
)
results.append(result)
session_index += 1
ok = sum(1 for r in results if "error" not in r) ok = sum(1 for r in results if "error" not in r)
logger.info("Done: %d/%d requests succeeded.", ok, len(results)) logger.info("Done: %d/%d requests succeeded.", ok, len(results))