GETs store the resolved page language, activity pings carry <html lang>
(including a new ping on in-place language switches, which also keeps the
switch's GET out of crawler classification), and trail items/crawler hits
surface it. The viewer shows flags discreetly: nothing on single-language
sites or primary-language-only rows, one leading flag for uniform visits,
transition flags on mixed trails, a flag row per crawler.
The stored ua_pretty froze each record at the uarite version of its
record time (old records showed disguised Meta crawlers as
"Chrome/145 Windows"). Client now carries a display-time-only
"uarite" field holding the full uarite.UA dataclass (pretty/engine/
os/provider/kind/url), filled when the viewer payload is built, so old
data always follows the current uarite. uarite 0.2.0 fixes the Meta
misdetection itself; _compact_user_agent is gone, tracking.py calls
uaparse directly.
Replace the pre-classified store (visits/crawlers/abuse lists written at
collection time, plus abuse_ips and in-memory pending/session tables) with
a raw append-only log: one Get record per document GET (full path, true
HTTP status, referer origin, preload flag) and one Msg per /_ws activity
message. Visitor/crawler/abuse classification, visit grouping (30-minute
inactivity gap), status/referer/UTM attribution and all aggregates are
derived in Store.display(), so future rule changes never invalidate stored
data. The viewer payload keeps its exact shape.
Fixes structurally:
- Abuser 404s on slug-format paths showed up as "articles read": the
not-found branch recorded the request twice through separate status
plumbing. Each request is now recorded once with its true status.
- 404 trail links never rendered red: cache-served navigations issue no
GET and the only real GET (the idle preload) was discarded before status
recording. Preloads are now recorded with pre=True, never counted, and
used for status attribution.
- formatAbuseRows merged a path's 404 probes and 200 reads into one entry;
the collapse is now keyed by (path, status class).
Rule improvements enabled by the redesign:
- The plain-404 abuse threshold counts within a 1-hour sliding window, so
long-time readers accumulating misses never classify (scanners spray).
- Hidden (admin) clients never trigger abuse classification: editing means
visiting not-found pages.
- /.well-known/ probes (RFC 8615, e.g. Chrome devtools) are never abuse
evidence; //foo-style empty path segments are an instant telltale.
- Visits can no longer open on an external exit URL; favicon fetches skip
hidden clients' referers/exits.
Legacy analytics.json files are set aside as .bak-legacy on startup.
Visits with under 5 s of total reported reading time are JS-running bots:
display() converts them to crawler hits (one per trail page, with referer
and UTM query) and excludes them from every aggregate. Crawler referers
join the favicon fetch origins and render with their icon in the crawler
table. Abuse hits now record the real response status, so the abuse table
splits 404 probes from the articles the abuser actually read (200 GETs),
shown as trail links like the visitor/crawler tables.
The first positional CLI argument (default localhost) names the site's
public hostname and its data directory under the cwd, replacing the
CWD-relative pagerite.* files. The public origin (https://<hostname>)
is now authoritative configuration instead of a value learned from
admin browsers: POST /_api/site-url, Data.site_url and the pagerite.js
reporter are removed, and page rendering, sitemap, robots.txt and
analytics own_origin all use SITE_URL consistently (localhost falls
back to the request's base URL).
Two analytics corrections verified against the production capture:
- pagerite.js suppressed pings with fr == '/_a', but fetch-navigation
away from the analytics page had already GET-ed the target without the
preload header; the orphaned pending hit then flushed to the crawler
list, classifying a real user as a crawler. Navigations away from /_a
now ping normally (the server rejects /_a as a target regardless, and
admin noise is already handled by hide=1).
- _remove_visit only reversed the visit's creation counts, leaving
views/transitions from later pings behind as orphans on the graph with
no matching row in the visitor table. An in-memory per-visit count log
now tracks every count event, so an admin hide=1 scrub reverses the
visit completely.
Document GETs now stash their status (200/404) in a pending table,
consumed by the matching ping: visits gain a per-path statuses map and
crawler hits a status field. Trail links with a 404 status render in
red with the status code in the tooltip, alongside the read time.
- Cull connections whose thin middle would render below ~0.8px
(MIN_WMID); drop external source/exit nodes whose connectors are
all culled, while site page nodes always stay
- Bead simulation persists across data reloads: emitters keyed per edge
direction, beads tracked by progress, so unrelated count changes no
longer reshuffle bead positions
- Bead speed relative to span length: constant 1.5s traversal per edge
- Top lane labeled with a house icon; all lane labels left-aligned just
past the source pill (half height on near-vertical branch lanes), with
guides running to the lane end so long slugs are never truncated
- weeklySeries shifts overlaid weeks onto the current week's time axis so
they overlay inside the plot instead of overflowing left; oldest weeks
paint first, current week on top
- Legend moved inside the visits chart's top right: current ISO week in
accent, past weeks as a single muted "Week M" / "Week M–N" specimen
- Past week curves use the muted color instead of faded accent
- Chart height reduced ~30% (180 -> 126)
- Charts render as single SVGs with axis labels inside the viewBox,
replacing the stretched plot + HTML overlay labels
- Rolling ranges end at now, t0 aligned to UTC day; bucket size follows
the window (6h up to 31 days) so "all" at its 30-day minimum renders
identically to "month"
- X labels always centered on their true position; no edge-align shifting
- rangeWindow simplified to rolling spans ending at now
- TransitionGraph "all" visual scale floored at the 30-day plot minimum
JS-running crawlers (Googlebot, GoogleOther, Applebot) execute pagerite.js
and send navigation pings, registering as visitors. Pings whose User-Agent
matches _is_bot_ua (any "bot" token plus listed exceptions) are now
ignored, so their document GETs flush to the crawler list as intended. No
source verification: a spoofed bot UA merely lands in the crawler stats,
and path-based abuse classification catches scanners regardless.
Idle-time link preloads from pagerite.js were queued as pending crawler
hits and flushed to the crawler list whenever the user navigated more than
10s later, so real visitors' subpage loads showed up as crawler hits.
Preload fetches now carry an x-pagerite-preload header and the document
GET handler skips tracking for them; the ping sent on actual navigation
does the counting.
Downloads the latest dbip-city-lite-YYYY-MM.mmdb.gz before starting the
server, skipping when the local database is current, falling back to the
previous month on 404, and removing older databases after an update.
Promotes httpx to a runtime dependency.
- keep visitor charts y-axis minimum range at 10
- keep 'all' chart x-axis minimum span at 30 days
- group crawler hits by (ip, ua) and list top pages visited, show crawler page load counts as N× prefix
- store and display geoip city, keep geoip country overwrite
- stream live updates over WebSocket /_api/ws/analytics
- include family ring arcs in transition map crop bounds
- remove top UA summary, limit crawlers to 10 and visits to 20
- human-readable relative timestamps with UTC tooltip
- Edge widths grow logarithmically with the connection count (~1 px at
a single count, uncapped); connections below 1% of total traffic are
pruned, bounding the graph to ~100 edges.
- Beads: per-direction flows emitted at a rate linear in the count,
each bead simulated independently in JS (no in-flight limit), offset
onto right-hand lanes so opposing flows don't collide, running under
the node circles with a glow.
- External links: referer origins as a node row above the map, exit
origins fanned outwards from their source page.
- Transitions are now stored per 5-minute bucket (sparse
from -> to -> bucket -> count) so the graph filters by time range
like the other series; legacy analytics files are discarded.
Add server-side visit analytics collection, a public-page ping endpoint,
and a full-screen AnalyticsView for admins.
Backend:
- Add pagerite/analytics.py: Analytics/Visit model, Store, and persistence
- Wire /_a ping endpoint and GET /_api/analytics into pagerite/app.py
Frontend:
- Add full-screen AnalyticsView with visitor charts and transition map
- Add VisitorCharts and TransitionGraph subcomponents
- Add analytics JS helpers in frontend/src/analytics/
- Send navigation pings from frontend/src/pagerite.js
- Mount AnalyticsView from frontend/src/main.js
- Document the feature in docs/analytics.md and update AGENTS.md