Compare commits

..
3 Commits
10 changed files with 276 additions and 171 deletions
+40 -20
View File
@@ -27,7 +27,8 @@ falsy values are omitted):
(an `fr` equal to `to` would log a bogus self-transition when a session (an `fr` equal to `to` would log a bogus self-transition when a session
already exists, e.g. a second tab). This ping is what starts already exists, e.g. a second tab). This ping is what starts
the visit and counts the entry page view — the document GET alone records the visit and counts the entry page view — the document GET alone records
nothing, so bots and admin browsing never register. JS-running crawlers nothing, so bots never register (admin browsing does register, but
flagged `hide`; see **Admins** below). JS-running crawlers
(Googlebot, GoogleOther, Applebot, ...) do ping, but their User-Agent (Googlebot, GoogleOther, Applebot, ...) do ping, but their User-Agent
gives them away: pings whose UA matches `_is_bot_ua` (anything calling gives them away: pings whose UA matches `_is_bot_ua` (anything calling
itself a "bot", plus known exceptions such as GoogleOther) are ignored itself a "bot", plus known exceptions such as GoogleOther) are ignored
@@ -57,12 +58,18 @@ falsy values are omitted):
without the preload header, and without the ping that GET would flush to without the preload header, and without the ping that GET would flush to
the crawler list. the crawler list.
- **Admins**: when SSO is in use and the session is known to be an admin, - **Admins**: when SSO is in use and the session is known to be an admin,
the client still pings but adds `hide=1`. The server then records the client still pings but adds `hide=1`. The activity is recorded as
nothing — and if the same client session already had a visit from before usual (navigations and all), but the `hide` flag is set on the **client
logging in, that visit is removed from the JSON along with every count record** — so it covers everything that client ever did: visits and
it recorded — an in-memory per-visit log of count events makes full crawler hits from before the login included. Hidden clients never appear
reversal possible. With no auth proxy (dev/test) in the viewer payload: `Store.display()` drops their visits, crawler
"admin" is everyone's state, so `hide` stays 0 and everything is recorded. hits, abuse hits and metadata, and computes every aggregate (site visits,
page views, transitions) from the visible visits only, so nothing needs
to be reversed or redacted. Pending crawler hits from a hidden client
are discarded when they expire, so admin browsing never lands in the
crawler list either. With no auth proxy
(dev/test) "admin" is everyone's state, so `hide` stays 0 and everything
is recorded.
- The server validates `to`: internal paths must be valid slug paths - The server validates `to`: internal paths must be valid slug paths
("/" or `[a-z0-9_-]` segments), external ones are re-derived to the ("/" or `[a-z0-9_-]` segments), external ones are re-derived to the
https origin and accepted only when the client sent exactly that. https origin and accepted only when the client sent exactly that.
@@ -94,13 +101,14 @@ falsy values are omitted):
the header only hides a GET from the crawler stats, the path-based abuse the header only hides a GET from the crawler stats, the path-based abuse
classification is unaffected). If a ping classification is unaffected). If a ping
from the same client arrives within 10 seconds the hit is discarded; from the same client arrives within 10 seconds the hit is discarded;
otherwise it is written to `crawlers`. Crawlers do not count as otherwise it is written to `crawlers` — unless the client is hidden
(admin), in which case the hit is discarded on expiry too. Crawlers do not count as
visits or views. The `Accept-Language` header is stored on the shared visits or views. The `Accept-Language` header is stored on the shared
`Client` immediately; reverse-DNS host names and DB-IP geoip `Client` immediately; reverse-DNS host names and DB-IP geoip
country/city are filled in asynchronously, just like for real visits. In country/city are filled in asynchronously, just like for real visits. In
the analytics viewer, crawler hits are grouped by client hash and shown as the analytics viewer, crawler hits are grouped by client hash and shown as
a trail of internal pages that crawler visited; the crawler table lists a trail of internal pages that crawler visited; the crawler table lists
the most active crawlers first rather than the most recent hits. the most recent crawler first, with the most active as a tie-breaker.
- **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment - **Abuse (scanner) hits**: a 404 for a telltale path — any URL segment
starting with a dot (`/.env`, `/.git/config`) or ending in `.php` starting with a dot (`/.env`, `/.git/config`) or ending in `.php`
classifies the source IP as abuse immediately, and ten plain 404s from one classifies the source IP as abuse immediately, and ten plain 404s from one
@@ -141,7 +149,10 @@ Each `Client` record:
- `city` — city name from the DB-IP MMDB lookup, when available, - `city` — city name from the DB-IP MMDB lookup, when available,
- `ua` — raw `User-Agent` string, - `ua` — raw `User-Agent` string,
- `ua_pretty` — compact display form of the UA (browser/OS/device) when - `ua_pretty` — compact display form of the UA (browser/OS/device) when
parsable, otherwise the raw string. parsable, otherwise the raw string,
- `hide` — true for admin clients (`hide=1` ping): all their visits,
crawler hits and abuse hits are recorded but excluded from every
statistic and from the viewer payload.
Each `Visit` record: Each `Visit` record:
@@ -149,13 +160,16 @@ Each `Visit` record:
- `entry` — first page (path) seen, - `entry` — first page (path) seen,
- `referer` — external https origin of the initial load, `""` for direct, - `referer` — external https origin of the initial load, `""` for direct,
- `client` — 6-byte blake3 hash referencing `Analytics.clients`, - `client` — 6-byte blake3 hash referencing `Analytics.clients`,
- `trail` — everything seen afterwards in first-seen order: page paths and - `trail` the entry page and everything seen afterwards, keyed by the
external exit URLs. Re-visiting an already seen page (incl. the entry) timestamp of first sight (insertion order = first-seen order). Each item
does not append. holds `to` (page path or external exit URL), the accumulated active
reading time in seconds (`read`) and the most recent HTTP status seen
for the target (`status`). Re-visiting an already seen target updates
its item instead of appending.
- `navs` — every navigation ping (`fr`, `to`), keyed by its timestamp,
repeats included. The aggregates are computed from this log at display
time.
- `utm``utm_*` query parameters from the landing URL, as a dict. - `utm``utm_*` query parameters from the landing URL, as a dict.
- `read` — active reading time per path (seconds), keyed by path.
- `statuses` — HTTP status of the response when each path was first seen
(200 or 404), keyed by path.
Each `CrawlerHit` record: Each `CrawlerHit` record:
@@ -190,10 +204,16 @@ to tell misses from real pages at a glance.
## Aggregates ## Aggregates
Aggregates are **not stored**; they are computed at display time by
`Store.display()` from the visit records (entry + `navs` log), skipping
hidden clients' visits. This is what allows a client to become hidden after
navigations were already logged: no counts need reversing. The computed
shapes, part of the WebSocket payload (`Display` struct alongside `visits`,
`crawlers`, `abuse` and `clients`):
- `transitions`: time series of page transitions, sparse nested dict - `transitions`: time series of page transitions, sparse nested dict
`from -> to -> bucket -> count` with the same 5-minute bucketing as `from -> to -> bucket -> count` with 5-minute bucketing. `from` is the
`views`. `from` is the referer origin or `"(direct)"` for initial loads, referer origin or `"(direct)"` for initial loads, a page path for pings.
a page path for pings.
- `views`: time series of page loads, `path -> bucket -> count`, sparse: only - `views`: time series of page loads, `path -> bucket -> count`, sparse: only
non-zero 5-minute buckets exist (bucket key is its floored ISO timestamp). non-zero 5-minute buckets exist (bucket key is its floored ISO timestamp).
Every load counts, including repeats within a visit; external exit origins Every load counts, including repeats within a visit; external exit origins
@@ -202,7 +222,7 @@ to tell misses from real pages at a glance.
5-minute bucketing. 5-minute bucketing.
Sparseness keeps quiet sites small; dropping old data is a matter of deleting Sparseness keeps quiet sites small; dropping old data is a matter of deleting
list/dict entries (`visits` is a plain append-only list, buckets plain keys). list entries (`visits` is a plain append-only list).
## Persistence ## Persistence
+21 -13
View File
@@ -80,17 +80,27 @@ export function calcTotalViews(views) {
// Very short reads are navigation/skims, not real reading time. // Very short reads are navigation/skims, not real reading time.
export const MIN_READ_SECONDS = 10 export const MIN_READ_SECONDS = 10
/** path -> accumulated read seconds for a visit, derived from its trail. */
export function readMapOf(v) {
const map = {}
for (const item of Object.values(v.trail || {})) {
if (item.read) map[item.to] = (map[item.to] || 0) + item.read
}
return map
}
/** Average minutes per visit and average of per-article median read minutes. */ /** Average minutes per visit and average of per-article median read minutes. */
export function calcReadStats(visits) { export function calcReadStats(visits) {
const perArticle = {} const perArticle = {}
let totalVisitSeconds = 0 let totalVisitSeconds = 0
let visitCount = 0 let visitCount = 0
for (const v of visits || []) { for (const v of visits || []) {
const secs = Object.values(v.read || {}).filter((s) => s >= MIN_READ_SECONDS) const read = readMapOf(v)
const secs = Object.values(read).filter((s) => s >= MIN_READ_SECONDS)
if (!secs.length) continue if (!secs.length) continue
visitCount++ visitCount++
totalVisitSeconds += secs.reduce((a, b) => a + b, 0) totalVisitSeconds += secs.reduce((a, b) => a + b, 0)
for (const [path, s] of Object.entries(v.read || {})) { for (const [path, s] of Object.entries(read)) {
if (s >= MIN_READ_SECONDS) { if (s >= MIN_READ_SECONDS) {
; (perArticle[path] || (perArticle[path] = [])).push(s) ; (perArticle[path] || (perArticle[path] = [])).push(s)
} }
@@ -271,7 +281,7 @@ export function formatRecentVisits(visits, pageTree, limit = 50) {
.reverse() .reverse()
.map((v) => ({ .map((v) => ({
when: new Date(v.start).toLocaleString(), when: new Date(v.start).toLocaleString(),
steps: [v.referer, v.entry, ...(v.trail || [])] steps: [v.referer, ...Object.values(v.trail || {}).map((t) => t.to)]
.map((p) => stepOf(p, titles)) .map((p) => stepOf(p, titles))
.filter(Boolean), .filter(Boolean),
})) }))
@@ -349,8 +359,8 @@ export function mainDomain(host, limit = 24) {
/** /**
* Group raw crawler hits by client hash and format each group as a row showing * Group raw crawler hits by client hash and format each group as a row showing
* every internal page that crawler visited. Rows are sorted by total hits, * every internal page that crawler visited. Rows are sorted by most recent hit
* most active crawler first, rather than by most recent hit. * first, with total hits as a tie-breaker.
* ``clients`` maps client hashes to client records. * ``clients`` maps client hashes to client records.
*/ */
export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) { export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now()) {
@@ -380,7 +390,7 @@ export function formatCrawlerRows(crawlers, clients, pageTree, now = Date.now())
return n return n
} }
return [...groups.values()] return [...groups.values()]
.sort((a, b) => totalHits(b) - totalHits(a) || b.lastStart - a.lastStart) .sort((a, b) => b.lastStart - a.lastStart || totalHits(b) - totalHits(a))
.slice(0, 10) .slice(0, 10)
.map((g) => { .map((g) => {
const client = g.client || {} const client = g.client || {}
@@ -510,14 +520,12 @@ export function formatVisitRows(visits, clients, pageTree, now = Date.now()) {
const titles = buildTitleMap(pageTree) const titles = buildTitleMap(pageTree)
return [...(visits || [])].reverse().slice(0, 20).map((v) => { return [...(visits || [])].reverse().slice(0, 20).map((v) => {
const client = (clients || {})[v.client] || {} const client = (clients || {})[v.client] || {}
const read = v.read || {} const trail = Object.values(v.trail || {})
const statuses = v.statuses || {} .map((item) => {
const trail = [v.entry, ...(v.trail || [])] const step = stepOf(item.to, titles)
.map((p) => {
const step = stepOf(p, titles)
if (step) { if (step) {
if (read[p]) step.readSeconds = read[p] if (item.read) step.readSeconds = item.read
if (statuses[p]) step.status = statuses[p] if (item.status) step.status = item.status
} }
return step return step
}) })
+2 -2
View File
@@ -29,7 +29,7 @@
* pings) are skipped. * pings) are skipped.
*/ */
import { MIN_READ_SECONDS } from './format.js' import { MIN_READ_SECONDS, readMapOf } from './format.js'
// Nodes are constant-size pills (stadium rects) holding the slug and the // Nodes are constant-size pills (stadium rects) holding the slug and the
// view count on two centered lines. TNODE_BOUND is the pill's bounding // view count on two centered lines. TNODE_BOUND is the pill's bounding
@@ -263,7 +263,7 @@ function sortByNav(root, navOrder) {
function buildReadSeconds(visits) { function buildReadSeconds(visits) {
const times = {} const times = {}
for (const v of visits || []) { for (const v of visits || []) {
for (const [path, sec] of Object.entries(v.read || {})) { for (const [path, sec] of Object.entries(readMapOf(v))) {
if (sec >= MIN_READ_SECONDS) { if (sec >= MIN_READ_SECONDS) {
; (times[path] || (times[path] = [])).push(sec) ; (times[path] || (times[path] = [])).push(sec)
} }
+7 -7
View File
@@ -928,10 +928,10 @@ figure:has(img[width]) {
later rules win at equal specificity. The analytics dashboard uses the later rules win at equal specificity. The analytics dashboard uses the
same breakout directly on its container (div.wide — it is the page's same breakout directly on its container (div.wide — it is the page's
whole content, not a figure), and code blocks via a trailing {.wide} whole content, not a figure), and code blocks via a trailing {.wide}
line (fence attrs land on <code>, hence pre:has(.wide)). */ line (fence block attrs land on <pre> itself). */
figure:has(.wide), figure:has(.wide),
div.wide, div.wide,
pre:has(.wide) { pre.wide {
width: 100vw; width: 100vw;
max-width: none; max-width: none;
margin-inline: calc(50% - 50vw); margin-inline: calc(50% - 50vw);
@@ -939,7 +939,7 @@ pre:has(.wide) {
/* Full bleed means edge to edge — no rounded corners. */ /* Full bleed means edge to edge — no rounded corners. */
figure:has(.wide) img, figure:has(.wide) img,
pre:has(.wide) { pre.wide {
border-radius: 0; border-radius: 0;
} }
@@ -948,7 +948,7 @@ pre:has(.wide) {
1fr + 4fr grid, i.e. 20vw — plus main's padding, and spans on to the 1fr + 4fr grid, i.e. 20vw — plus main's padding, and spans on to the
right viewport edge. */ right viewport edge. */
body:has(.multicol) figure:has(.wide), body:has(.multicol) figure:has(.wide),
body:has(.multicol) pre:has(.wide) { body:has(.multicol) pre.wide {
margin-inline: calc(-20vw - 1.25rem) 0; margin-inline: calc(-20vw - 1.25rem) 0;
} }
@@ -956,7 +956,7 @@ body:has(.multicol) pre:has(.wide) {
window keeps its overlay scrollbars while editing, so — unlike a classic window keeps its overlay scrollbars while editing, so — unlike a classic
scrollbar — they take no layout space and the vw math stays exact. */ scrollbar — they take no layout space and the vw math stays exact. */
body.editing figure:has(.wide), body.editing figure:has(.wide),
body.editing pre:has(.wide) { body.editing pre.wide {
width: calc(100vw - var(--editor-w)); width: calc(100vw - var(--editor-w));
margin-inline: calc(50% - (100vw - var(--editor-w)) / 2); margin-inline: calc(50% - (100vw - var(--editor-w)) / 2);
} }
@@ -964,7 +964,7 @@ body.editing pre:has(.wide) {
/* Editing + multicol: the left gutter is 1/5 of the space right of the /* Editing + multicol: the left gutter is 1/5 of the space right of the
editor, and the bleed also crosses main's 1.25rem left padding. */ editor, and the bleed also crosses main's 1.25rem left padding. */
body.editing:has(.multicol) figure:has(.wide), body.editing:has(.multicol) figure:has(.wide),
body.editing:has(.multicol) pre:has(.wide) { body.editing:has(.multicol) pre.wide {
margin-inline: calc((100vw - var(--editor-w)) / -5 - 1.25rem) 0; margin-inline: calc((100vw - var(--editor-w)) / -5 - 1.25rem) 0;
} }
@@ -977,7 +977,7 @@ body.editing:has(.multicol) pre:has(.wide) {
where the editing rules above apply instead. */ where the editing rules above apply instead. */
@media (max-width: 102rem) { @media (max-width: 102rem) {
body:has(#sidebar):not(.editing) figure:has(.wide), body:has(#sidebar):not(.editing) figure:has(.wide),
body:has(#sidebar):not(.editing) pre:has(.wide) { body:has(#sidebar):not(.editing) pre.wide {
margin-inline: -13.25rem 0; margin-inline: -13.25rem 0;
} }
} }
+20 -3
View File
@@ -235,9 +235,26 @@ import "overlayscrollbars/overlayscrollbars.css";
btn.textContent = "copy"; btn.textContent = "copy";
btn.addEventListener("click", async () => { btn.addEventListener("click", async () => {
const code = pre.querySelector("code"); const code = pre.querySelector("code");
await navigator.clipboard.writeText( const text = (code || pre).textContent.replace(/\n$/, "");
(code || pre).textContent.replace(/\n$/, ""), // navigator.clipboard exists only in secure contexts (https or
); // localhost); viewing over plain http needs the textarea fallback.
try {
if (navigator.clipboard) {
await navigator.clipboard.writeText(text);
} else {
const ta = document.createElement("textarea");
ta.value = text;
ta.style.cssText = "position:fixed;opacity:0";
document.body.append(ta);
ta.select();
document.execCommand("copy");
ta.remove();
}
} catch {
btn.textContent = "failed";
setTimeout(() => (btn.textContent = "copy"), 1500);
return;
}
btn.textContent = "copied"; btn.textContent = "copied";
btn.classList.add("copied"); btn.classList.add("copied");
setTimeout(() => { setTimeout(() => {
+155 -116
View File
@@ -12,8 +12,13 @@ away (``_is_bot_ua``) and their pings are ignored, so they land in the
crawler list too. Idle-time link preloads from pagerite.js carry an crawler list too. Idle-time link preloads from pagerite.js carry an
``x-pagerite-preload`` header and are not tracked at all — the ping sent ``x-pagerite-preload`` header and are not tracked at all — the ping sent
when the user actually navigates does the counting. when the user actually navigates does the counting.
Admin clients ping with ``hide=1``, which records nothing and removes any Admin clients ping with ``hide=1``: the client record is flagged ``hide``,
visit the session accumulated before logging in. Scanner telltale 404s which covers everything that client ever did — visits and crawler hits
from before the login included. Aggregates (site visits, page views,
transitions) are not stored; they are computed at display time from the
visit records, excluding hidden clients, and hidden clients' visits,
crawler hits, abuse hits and metadata are left out of the viewer payload
entirely. Scanner telltale 404s
(dotpaths, *.php) classify the source IP as abuse; its hits — including (dotpaths, *.php) classify the source IP as abuse; its hits — including
earlier crawler hits — are moved to the abuse list, which the viewer earlier crawler hits — are moved to the abuse list, which the viewer
groups by IP with full request paths. Client metadata (IP, UA, language, groups by IP with full request paths. Client metadata (IP, UA, language,
@@ -87,15 +92,47 @@ class Client(msgspec.Struct, omit_defaults=True):
ua: str = "" ua: str = ""
#: Compact display form of ``ua`` (browser/OS/device) when parsable. #: Compact display form of ``ua`` (browser/OS/device) when parsable.
ua_pretty: str = "" ua_pretty: str = ""
#: True for admin clients (hide=1 ping): their visits, crawler hits and
#: abuse hits are recorded but excluded from all statistics and from
#: the viewer payload.
hide: bool = False
class Nav(msgspec.Struct, omit_defaults=True):
"""One navigation inside a visit: from ``fr`` to ``to``.
``to`` is an internal page path or an external https exit URL. Every
navigation is logged (repeats included), keyed by its timestamp in
``Visit.navs``, so display-time aggregates can count views and
transitions; ``Visit.trail`` keeps the first-seen order.
"""
fr: str
to: str
class TrailItem(msgspec.Struct, omit_defaults=True):
"""One first-seen target in a visit trail: a page or external exit URL.
``read`` accumulates active reading time (seconds) across the whole
visit; ``status`` is the most recent HTTP status seen for the target.
"""
to: str
#: Accumulated active reading time in seconds.
read: int = 0
#: Most recent HTTP status of the response (200 or 404).
status: int = 200
class Visit(msgspec.Struct, omit_defaults=True): class Visit(msgspec.Struct, omit_defaults=True):
"""One visit: the initial-load data plus everything seen afterwards. """One visit: the initial-load data plus everything seen afterwards.
``trail`` holds page paths and external exit URLs in first-seen ``trail`` holds the entry page and everything seen afterwards, keyed by
order; re-visiting an already seen page does not append. The entry the timestamp of first sight (insertion order = first-seen order);
page itself is in ``entry``, not in the trail. Client metadata is re-visiting an already seen target updates its item instead of
held in ``Analytics.clients`` keyed by ``client``. appending. Client metadata is held in ``Analytics.clients`` keyed by
``client``.
""" """
start: datetime start: datetime
@@ -104,13 +141,13 @@ class Visit(msgspec.Struct, omit_defaults=True):
referer: str = "" referer: str = ""
#: 6-byte blake3 hash referencing ``Analytics.clients``. #: 6-byte blake3 hash referencing ``Analytics.clients``.
client: bytes = b"" client: bytes = b""
trail: list[str] = [] #: First-seen targets keyed by their timestamp (entry included).
trail: dict[datetime, TrailItem] = {}
#: Every navigation ping (repeats included) keyed by its timestamp; the
#: aggregates are computed from this log at display time.
navs: dict[datetime, Nav] = {}
#: UTM query parameters from the landing URL, keyed by parameter name. #: UTM query parameters from the landing URL, keyed by parameter name.
utm: dict[str, str] = {} utm: dict[str, str] = {}
#: Active reading time per path (seconds), keyed by path.
read: dict[str, int] = {}
#: HTTP status of the response when the path was first seen (200 or 404).
statuses: dict[str, int] = {}
class CrawlerHit(msgspec.Struct, omit_defaults=True): class CrawlerHit(msgspec.Struct, omit_defaults=True):
@@ -166,6 +203,22 @@ class Analytics(msgspec.Struct, omit_defaults=True):
clients: dict[bytes, Client] = {} clients: dict[bytes, Client] = {}
#: IPs classified as scanners/abusers (keys; values always True). #: IPs classified as scanners/abusers (keys; values always True).
abuse_ips: dict[str, bool] = {} abuse_ips: dict[str, bool] = {}
class Display(msgspec.Struct, omit_defaults=True):
"""The viewer payload: visible data plus display-time aggregates.
Hidden clients are excluded everywhere: their visits, crawler hits,
abuse hits and metadata are dropped, and the aggregates are computed
from the visible visits only.
The aggregate shapes match what the viewer consumes: sparse 5-minute
buckets keyed by their floored ISO timestamp.
"""
visits: list[Visit] = []
crawlers: list[CrawlerHit] = []
abuse: list[AbuseHit] = []
clients: dict[bytes, Client] = {}
#: Page transitions per 5-minute bucket (sparse): #: Page transitions per 5-minute bucket (sparse):
#: from -> to -> bucket ISO -> count. ``from`` is the referer origin or #: from -> to -> bucket ISO -> count. ``from`` is the referer origin or
#: "(direct)" for initial loads, a page path for pings. #: "(direct)" for initial loads, a page path for pings.
@@ -316,10 +369,6 @@ class Store:
pass # legacy schema / corrupt or unreadable file: start fresh pass # legacy schema / corrupt or unreadable file: start fresh
#: client hash -> index of the current visit in data.visits #: client hash -> index of the current visit in data.visits
self.sessions: dict[bytes, int] = {} self.sessions: dict[bytes, int] = {}
#: visit index -> count events recorded for that visit, so
#: ``_remove_visit`` can reverse all of them — not just the ones
#: from the visit's creation. In-memory only, like ``sessions``.
self._count_log: dict[int, list[tuple]] = {}
#: ip -> external https origin of the latest document GET carrying #: ip -> external https origin of the latest document GET carrying
#: one, stashed for the visit the client's initial ping starts. #: one, stashed for the visit the client's initial ping starts.
#: Internal or absent referers never touch the table. #: Internal or absent referers never touch the table.
@@ -373,6 +422,9 @@ class Store:
def _flush_crawlers(self, now: datetime | None = None) -> list[bytes]: def _flush_crawlers(self, now: datetime | None = None) -> list[bytes]:
"""Move expired pending crawler hits into persistent ``data.crawlers``. """Move expired pending crawler hits into persistent ``data.crawlers``.
Hits from a hidden client (admin) are discarded instead of
persisted — admin browsing must not land in the crawler list.
Returns the client hashes of the newly flushed hits so callers can Returns the client hashes of the newly flushed hits so callers can
schedule async enrichment. schedule async enrichment.
""" """
@@ -383,68 +435,63 @@ class Store:
expired: list[CrawlerHit] = [] expired: list[CrawlerHit] = []
remaining: list[CrawlerHit] = [] remaining: list[CrawlerHit] = []
for hit in self.pending_crawlers: for hit in self.pending_crawlers:
(expired if hit.start <= cutoff else remaining).append(hit) if hit.start > cutoff:
remaining.append(hit)
continue
client = self.data.clients.get(hit.client)
if client is not None and client.hide:
continue # hidden admin client: not a crawler
expired.append(hit)
if not expired: if not expired:
self.pending_crawlers = remaining
return [] return []
self.pending_crawlers = remaining self.pending_crawlers = remaining
self.data.crawlers.extend(expired) self.data.crawlers.extend(expired)
self._save() self._save()
return [hit.client for hit in expired] return [hit.client for hit in expired]
def _count(self, table: dict[str, int], key: str) -> None: def _hidden(self, client_hash: bytes) -> bool:
table[key] = table.get(key, 0) + 1 """True when the client record is flagged hidden (admin)."""
client = self.data.clients.get(client_hash)
return client is not None and client.hide
def _count_transition(self, fr: str, to: str, now: datetime) -> None: def display(self) -> Display:
"""Count one transition in its 5-minute bucket (sparse matrix).""" """Build the viewer payload, excluding hidden clients.
buckets = self.data.transitions.setdefault(fr, {}).setdefault(to, {})
self._count(buckets, _bucket(now))
def _uncount(self, table: dict[str, int], key: str) -> None: The aggregates (site visits, page views, transitions) are computed
"""Reverse one ``_count``: decrement and drop empty keys.""" here from the visit records rather than stored, so a client that
if key in table: becomes hidden after navigations were already logged disappears
table[key] -= 1 from every statistic. Internal-path navigations count as page
if table[key] <= 0: views; external https targets are transitions only.
del table[key]
def _remove_visit(self, index: int) -> None:
"""Delete a visit and reverse every count it recorded.
Used when a known visitor turns out to be an admin (hide=1 ping):
the session is scrubbed from the stats. The in-memory
``_count_log`` tracks each site-visit/view/transition count the
visit produced, so the scrub reverses all of them — including the
ones logged by later pings inside the visit.
""" """
for event in self._count_log.pop(index, ()): visits = [v for v in self.data.visits if not self._hidden(v.client)]
kind = event[0] display = Display(
if kind == "site": visits=visits,
self._uncount(self.data.site_visits, event[1]) crawlers=[h for h in self.data.crawlers if not self._hidden(h.client)],
elif kind == "view": abuse=[h for h in self.data.abuse if not self._hidden(h.client)],
views = self.data.views.get(event[1]) clients={h: c for h, c in self.data.clients.items() if not c.hide},
if views is not None: )
self._uncount(views, event[2]) for visit in visits:
if not views: bucket = _bucket(visit.start)
del self.data.views[event[1]] site = display.site_visits
else: # transition site[bucket] = site.get(bucket, 0) + 1
_, fr, to, bucket = event entry_views = display.views.setdefault(visit.entry, {})
fr_map = self.data.transitions.get(fr) entry_views[bucket] = entry_views.get(bucket, 0) + 1
if fr_map is not None: fr = visit.referer or "(direct)"
buckets = fr_map.get(to) buckets = display.transitions.setdefault(fr, {}).setdefault(visit.entry, {})
if buckets is not None: buckets[bucket] = buckets.get(bucket, 0) + 1
self._uncount(buckets, bucket) for t, nav in visit.navs.items():
if not buckets: nb = _bucket(t)
del fr_map[to] if nav.to.startswith("/"):
if not fr_map: nav_views = display.views.setdefault(nav.to, {})
del self.data.transitions[fr] nav_views[nb] = nav_views.get(nb, 0) + 1
del self.data.visits[index] nbuckets = display.transitions.setdefault(nav.fr, {}).setdefault(nav.to, {})
# Sessions and count logs store list indices; shift the ones past nbuckets[nb] = nbuckets.get(nb, 0) + 1
# the removed visit. return display
for key, i in list(self.sessions.items()):
if i > index: def display_json(self) -> str:
self.sessions[key] = i - 1 """The ``display()`` payload as a JSON string for the WebSocket."""
self._count_log = { return msgspec.json.encode(self.display()).decode()
i - 1 if i > index else i: log for i, log in self._count_log.items()
}
def _client_ip(self, client_hash: bytes) -> str: def _client_ip(self, client_hash: bytes) -> str:
"""Return the IP stored for ``client_hash``, or "" if missing.""" """Return the IP stored for ``client_hash``, or "" if missing."""
@@ -601,20 +648,9 @@ class Store:
client=client_hash, client=client_hash,
utm=utm or {}, utm=utm or {},
) )
visit.statuses[entry] = status visit.trail[now] = TrailItem(to=entry, status=status)
self.data.visits.append(visit) self.data.visits.append(visit)
index = len(self.data.visits) - 1 self.sessions[client_hash] = len(self.data.visits) - 1
self.sessions[client_hash] = index
bucket = _bucket(now)
fr = referer or "(direct)"
self._count(self.data.site_visits, bucket)
self._count(self.data.views.setdefault(entry, {}), bucket)
self._count_transition(fr, entry, now)
self._count_log[index] = [
("site", bucket),
("view", entry, bucket),
("transition", fr, entry, bucket),
]
return visit return visit
def track_entry( def track_entry(
@@ -688,7 +724,10 @@ class Store:
if index is None or index >= len(self.data.visits): if index is None or index >= len(self.data.visits):
return return
visit = self.data.visits[index] visit = self.data.visits[index]
visit.read[path] = visit.read.get(path, 0) + seconds for item in visit.trail.values():
if item.to == path:
item.read += seconds
return
def ping( def ping(
self, self,
@@ -711,10 +750,11 @@ class Store:
A ping with no known session starts a fresh visit, consuming the A ping with no known session starts a fresh visit, consuming the
referer and UTM tags stashed by the document GET if there are any. referer and UTM tags stashed by the document GET if there are any.
``hide`` is set by admin clients: the ping cancels pending crawler ``hide`` is set by admin clients: the client record is flagged
hits as usual, and any existing visit for this client session is ``hide`` — which covers everything it ever did, including visits and
removed from the stats (the admin browsed anonymously before logging crawler hits from before the login — and the navigation is recorded
in). Nothing new is recorded. normally. Hidden clients are excluded from every statistic and list
at display time, and their pending crawler hits are discarded.
Pings from IPs classified as abuse, and pings whose User-Agent Pings from IPs classified as abuse, and pings whose User-Agent
claims a JS-running crawler identity (``_is_bot_ua``), are ignored claims a JS-running crawler identity (``_is_bot_ua``), are ignored
@@ -727,34 +767,35 @@ class Store:
""" """
flushed = self._flush_crawlers() flushed = self._flush_crawlers()
lang, country = _parse_accept_language(accept_language) lang, country = _parse_accept_language(accept_language)
client_hash = _client_hash(ip, ua, lang)
if hide: if hide:
# Admin ping: cancel pending crawler hits and scrub the session. # Admin ping: flag the client hidden and never a crawler hit.
# The flag lives on the client record, so it covers visits and
# crawler hits from before the login too; display-time
# aggregation excludes hidden clients from every statistic.
client_hash = self._ensure_client(ip, ua, lang, country=country)
self.data.clients[client_hash].hide = True
self.pending_crawlers = [
hit for hit in self.pending_crawlers if hit.client != client_hash
]
else:
client_hash = _client_hash(ip, ua, lang)
if ip in self.data.abuse_ips:
return None, flushed
if _is_bot_ua(ua):
# A JS-running crawler (Googlebot, GoogleOther, Applebot
# execute JS and ping): never a visit. Its pending crawler
# hits are kept and flush to ``data.crawlers`` normally.
return None, flushed
# A real visitor ping cancels any pending crawler hits from
# this client.
self.pending_crawlers = [ self.pending_crawlers = [
hit for hit in self.pending_crawlers if hit.client != client_hash hit for hit in self.pending_crawlers if hit.client != client_hash
] ]
self.pending_statuses.pop(client_hash, None)
index = self.sessions.pop(client_hash, None)
if index is not None and index < len(self.data.visits):
self._remove_visit(index)
self._save()
return None, flushed
if ip in self.data.abuse_ips:
return None, flushed
if _is_bot_ua(ua):
# A JS-running crawler (Googlebot, GoogleOther, Applebot execute
# JS and ping): never a visit. Its pending crawler hits are
# kept and flush to ``data.crawlers`` normally.
return None, flushed
# A real visitor ping cancels any pending crawler hits from this client.
self.pending_crawlers = [
hit for hit in self.pending_crawlers if hit.client != client_hash
]
fr_path = _internal_path(from_) if from_ else "" fr_path = _internal_path(from_) if from_ else ""
if fr_path and read > 0: if fr_path and read > 0:
self._add_read(client_hash, fr_path, read) self._add_read(client_hash, fr_path, read)
if not to: if not to:
if read > 0: if read > 0 or hide:
self._save() self._save()
return None, flushed return None, flushed
if to.startswith("/") and not to.startswith("//"): if to.startswith("/") and not to.startswith("//"):
@@ -784,17 +825,15 @@ class Store:
else: else:
visit = self.data.visits[index] visit = self.data.visits[index]
now = datetime.now(UTC) now = datetime.now(UTC)
bucket = _bucket(now) visit.navs[now] = Nav(fr=fr, to=target)
log = self._count_log.setdefault(index, []) # First-seen only: repeat pages and repeated exits update the
if target.startswith("/"): # existing trail item (most recent status) instead of appending.
self._count(self.data.views.setdefault(target, {}), bucket) for item in visit.trail.values():
log.append(("view", target, bucket)) if item.to == target:
self._count_transition(fr, target, now) item.status = target_status
log.append(("transition", fr, target, bucket)) break
# First-seen only: repeat pages and repeated exits don't append. else:
if visit.entry != target and target not in visit.trail: visit.trail[now] = TrailItem(to=target, status=target_status)
visit.trail.append(target)
visit.statuses[target] = target_status
self._save() self._save()
visit_index = index if index is not None and index < len(self.data.visits) else None visit_index = index if index is not None and index < len(self.data.visits) else None
return visit_index, flushed return visit_index, flushed
+4 -3
View File
@@ -782,7 +782,7 @@ async def _broadcast_analytics() -> None:
"""Send the current analytics snapshot to every connected WS client.""" """Send the current analytics snapshot to every connected WS client."""
if not _analytics_ws_clients: if not _analytics_ws_clients:
return return
payload = msgspec.json.encode(analytics_store.data).decode() payload = analytics_store.display_json()
closed = set() closed = set()
for ws in _analytics_ws_clients: for ws in _analytics_ws_clients:
try: try:
@@ -865,7 +865,8 @@ def _track_entry(path: str, request: Request, *, status: int = 200) -> list[byte
Nothing is counted on the GET itself — the client's /_a ping starts the Nothing is counted on the GET itself — the client's /_a ping starts the
visit, so bots never register as visits (JS-running crawlers ping too, visit, so bots never register as visits (JS-running crawlers ping too,
but the ping handler ignores known bot UAs). (Admin clients ping too, but the ping handler ignores known bot UAs). (Admin clients ping too,
but with hide=1, which scrubs their session instead of recording it.) but with hide=1, which flags their visit hidden: it is recorded but
excluded from all statistics and from the crawler list.)
The devserver's health probe (``GET /?from=devserver.py`` from The devserver's health probe (``GET /?from=devserver.py`` from
``127.0.0.1``) is ignored: it is not real traffic and would otherwise be ``127.0.0.1``) is ignored: it is not real traffic and would otherwise be
@@ -941,7 +942,7 @@ async def analytics_websocket(ws: WebSocket) -> None:
endpoint. Powers the analytics viewer rendered at /_a. endpoint. Powers the analytics viewer rendered at /_a.
""" """
await ws.accept() await ws.accept()
await ws.send_text(msgspec.json.encode(analytics_store.data).decode()) await ws.send_text(analytics_store.display_json())
_analytics_ws_clients.add(ws) _analytics_ws_clients.add(ws)
try: try:
while True: while True:
+26
View File
@@ -78,6 +78,31 @@ def _highlight(text: str, lang: str, _attrs: str) -> str:
return highlight(text, lexer, _formatter) return highlight(text, lexer, _formatter)
def _fence_rule(
self: RendererHTML,
tokens,
idx: int,
options,
env: dict,
) -> str:
"""Render a fenced code block.
Like the default fence rule, but block attributes (a trailing `{...}`
line, applied to the fence token by _block_attrs) go on the <pre> — the
block element — instead of the <code>, which keeps only the language
class. This is what makes e.g. `{.wide}` or `{style="..."}` after a
code fence style the block itself.
"""
token = tokens[idx]
info = token.info.strip() if token.info else ""
lang = info.split(maxsplit=1)[0] if info else ""
highlighted = (_highlight(token.content, lang, "")
or escapeHtml(token.content))
code_class = f' class="{options.langPrefix}{lang}"' if lang else ""
return (f"<pre{self.renderAttrs(token)}><code{code_class}>"
f"{highlighted}</code></pre>\n")
def _image_rule( def _image_rule(
self: RendererHTML, self: RendererHTML,
tokens, tokens,
@@ -286,6 +311,7 @@ md = (
.use(superscript_plugin) .use(superscript_plugin)
) )
md.add_render_rule("image", _image_rule) md.add_render_rule("image", _image_rule)
md.add_render_rule("fence", _fence_rule)
# GFM alerts (`> [!NOTE]` etc.), built into markdown-it-py's blockquote rule. # GFM alerts (`> [!NOTE]` etc.), built into markdown-it-py's blockquote rule.
md.options["alerts"] = True md.options["alerts"] = True
# Block attrs must be stripped before the typographer curlifies their quotes. # Block attrs must be stripped before the typographer curlifies their quotes.
-1
View File
@@ -108,7 +108,6 @@ article h3 {
font-weight: 700; font-weight: 700;
font-size: 0.95rem; font-size: 0.95rem;
letter-spacing: 0.08em; letter-spacing: 0.08em;
text-transform: uppercase;
color: var(--muted); color: var(--muted);
} }
+1 -6
View File
@@ -138,14 +138,9 @@
/* Console-style headings: uppercase monospace. h1 in the page text color /* Console-style headings: uppercase monospace. h1 in the page text color
with a hazard-stripe underline, h2 deep orange, h3 cyan. */ with a hazard-stripe underline, h2 deep orange, h3 cyan. */
article h1, article h1 {
article h2,
article h3 {
text-transform: uppercase; text-transform: uppercase;
letter-spacing: 0.02em; letter-spacing: 0.02em;
}
article h1 {
color: var(--text); color: var(--text);
font-weight: 700; font-weight: 700;
padding-bottom: 0.5rem; padding-bottom: 0.5rem;