Rework bot taxonomy: provider field, social/analytics kinds, spoof detection

- Distinguish Meta's five documented crawlers (Facebook, FacebookBot,
  Meta-ExternalAgent/ExternalFetcher/WebIndexer) with distinct kinds
- Rename kind preview -> social for link-sharing unfurlers; BingPreview
  moves to search as it is a search-engine feature
- Add analytics kind for monitoring and ad/SEO measurement crawlers
  (UptimeRobot, Pingdom, MJ12bot, Mediapartners-Google, AdsBot-Google)
- Reclassify Google family by main product use: Storebot-Google and
  Feedfetcher-Google/InspectionTool are search, Read-Aloud and
  GoogleOther are ai; AhrefsBot is search
- Drop retired entries: DuplexWeb-Google, SkypeUriPreview, PhantomJS
- Detect the Qualys SSL Labs scanner by its frozen exact UA string
- Mark ancient (pre-2023) auto-updating browser claims with
  " (spoofed)" in pretty, leaving engine/os/kind empty
- UA dataclass: all fields default to "", new provider field,
  constructions use kwargs
- PROVIDERS is now dict[provider, frozenset[names]] with derived
  PROVIDER_OF and NAME_KIND reverse lookups
- README and prettytable.py updated to match
This commit is contained in:
2026-09-08 23:20:06 +00:00
parent b70c8f313c
commit 84b9e27ced
4 changed files with 130 additions and 75 deletions
+16 -12
View File
@@ -29,8 +29,9 @@ r.url # ""
r = uaparse("Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/134.0.6885.65 Mobile Safari/537.36; compatible; facebookexternalhit/1.1; +http://www.facebook.com/externalhit_uatext.php")
r.pretty # "Facebook"
r.kind # "preview"
r.kind # "social"
r.bot # "Facebook"
r.provider # "Meta"
r.url # "http://www.facebook.com/externalhit_uatext.php"
```
@@ -38,19 +39,22 @@ r.url # "http://www.facebook.com/externalhit_uatext.php"
`uaparse(ua)` returns a frozen `UA` dataclass:
| Field | Content |
| -------- | ------------------------------------------------------------------------------------------------ |
| `pretty` | Compact display string (below); `""` for empty/missing UAs, the raw UA when unrecognized |
| `engine` | `"Chromium"`, `"Gecko"`, `"Safari"`, `"ArkWeb"` (HarmonyOS), or `""` |
| `os` | `"Windows"`, `"macOS"`, `"Linux"`, `"iOS"`, `"Android"`, `"HarmonyOS"`, or `""` |
| `bot` | Crawler/previewer display name, or `""` |
| `kind` | `"browser"`, `"ai"`, `"search"`, `"preview"`, `"spider"`, or `""` (scripts/HTTP libraries) |
| `url` | The crawler's info URL (`+https://…` pointer), or `""`; not part of `pretty` — link it in the UI |
| Field | Content |
| -------- | ------------------------------------------------------------------------------------------------- |
| pretty | Compact display string (below); empty for empty/missing UAs, the raw UA when unrecognized |
| engine | Chromium, Gecko, Safari, ArkWeb (HarmonyOS), or empty |
| os | Windows, macOS, Linux, iOS, Android, HarmonyOS, or empty |
| bot | Crawler/unfurler display name, or empty |
| kind | browser, ai, search, social, analytics, spider, or empty (scripts/HTTP libraries) |
| url | The crawler's info URL (the +https://… pointer), or empty; not part of pretty — link it in the UI |
| provider | The bot's provider for known crawler families (Meta, Google, OpenAI, ...), or empty |
`os` is the major OS only, no version — meant for things like offering OS-specific downloads. `engine` is derived from the browser identity: every recognized browser is Chromium except Firefox/LibreWolf (Gecko) and Safari and all of iOS (Safari's engine is all Apple allows there); HarmonyOS browsers run ArkWeb. Both are left empty for crawlers: the browser and OS in a disguised crawler UA are part of the disguise.
`kind` is `"browser"` for Mozilla-format UAs with no bot token, `"ai"` for training-data and AI-assistant fetchers (GPTBot, ClaudeBot, Google-Extended, ...), `"search"` for search-engine indexing (Googlebot, Bingbot, ...), `"preview"` for social link-preview fetchers (Facebook, WhatsApp, Slack, ...), `"spider"` for generic or unknown crawlers, and `""` for scripts and HTTP libraries.
The kind field describes our detection of visitor type: browser for actual browsers, ai for AI training collectors, agents and user-initiated fetches (GPTBot, ClaudeBot, ChatGPT-User, Google-Extended, ...), search for search-engine indexing (Googlebot, Bingbot, ...), social for link-sharing unfurlers (Facebook, WhatsApp, Slack, ...), analytics for monitoring and site-analytics crawlers (UptimeRobot, AdsBot-Google, MJ12bot), spider for generic or unknown crawlers, and empty for scripts and HTTP libraries. Any value other than browser means the visitor is automated.
`pretty` is intended to be shown directly:
- Desktop: `Chrome/152 Windows`, `Safari/18 macOS`
@@ -90,11 +94,11 @@ The table below compares representative results. uarite shows `r.pretty`; the ua
| Huawei HarmonyOS phone | HuaweiBrowser/6 HarmonyOS | Huawei Browser/6 Android❌ Huawei Browser |
| GPTBot | GPTBot (AI) | GPTBot/1 Spider |
| Googlebot (disguised) | Googlebot (search) | Googlebot/2 Android❌ Spider |
| Facebook preview (disguised) | Facebook | FacebookBot/1 Android Pixel 7 ❌ |
| Meta crawler (disguised) | Meta | Chrome/145 Windows ❌ |
| Facebook preview (disguised) | Facebook (social) | FacebookBot/1 Android Pixel 7 ❌ |
| Meta crawler (disguised) | Meta-ExternalAgent (AI) | Chrome/145 Windows ❌ |
| WhatsApp preview | WhatsApp | WhatsApp/10 Spider |
| Bytespider | Bytespider | Bytespider/ Android❌ Generic Smartphone |
| BingPreview | BingPreview (preview) | BingPreview/1 Windows❌ Spider |
| BingPreview | BingPreview | BingPreview/1 Windows❌ Spider |
| AhrefsBot | AhrefsBot | AhrefsBot/7 Spider |
| python-requests | python-requests/2.32.5 | Python Requests/2 |
-4
View File
@@ -55,10 +55,6 @@ CASES = [
"Googlebot (disguised)",
"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.84 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
),
(
"Claude-SearchBot (disguised)",
"Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/148.0.3245.171 Mobile Safari/537.36; compatible; Claude-SearchBot/1.0; +https://www.anthropic.com/claude-searchbot",
),
(
"Facebook preview (disguised)",
"Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/134.0.6885.65 Mobile Safari/537.36; compatible; facebookexternalhit/1.1; +http://www.facebook.com/externalhit_uatext.php",
+67 -42
View File
@@ -1,19 +1,24 @@
"""Crawler and link-preview tables."""
"""Crawler and link-unfurler tables."""
#: Lowercase UA substring to (display name, kind). Ordered: first match
#: wins, so overlapping names go from most to least specific.
BOTS = {
# Qualys SSL Labs scanner: frozen on this exact Firefox/45 string since
# ~2016; the pinned Gecko date makes the substring distinctive.
"mozilla/5.0 (x11; linux x86_64; rv:45.0) gecko/20100101 firefox/45.0": (
"Qualys SSL Labs",
"spider",
),
"feedfetcher-google": ("Feedfetcher-Google", "search"),
"google-inspectiontool": ("Google-InspectionTool", "search"),
"google-read-aloud": ("Google-Read-Aloud", "search"),
"mediapartners-google": ("Mediapartners-Google", "spider"),
"adsbot-google": ("AdsBot-Google", "spider"),
"google-read-aloud": ("Google-Read-Aloud", "ai"),
"mediapartners-google": ("Mediapartners-Google", "analytics"),
"adsbot-google": ("AdsBot-Google", "analytics"),
"apis-google": ("APIs-Google", "spider"),
"storebot-google": ("Storebot-Google", "spider"),
"duplexweb-google": ("DuplexWeb-Google", "spider"),
"storebot-google": ("Storebot-Google", "search"),
"google-extended": ("Google-Extended", "ai"),
"googlebot": ("Googlebot", "search"),
"googleother": ("GoogleOther", "spider"),
"googleother": ("GoogleOther", "ai"),
"bingbot": ("Bingbot", "search"),
"applebot": ("Applebot", "search"),
"gptbot": ("GPTBot", "ai"),
@@ -29,34 +34,40 @@ BOTS = {
"reflectionbot": ("Reflectionbot", "ai"),
"amzn-searchbot": ("Amzn-SearchBot", "search"),
"amazonbot": ("Amazonbot", "search"),
"ahrefsbot": ("AhrefsBot", "spider"),
"mj12bot": ("MJ12bot", "spider"),
"facebookexternalhit": ("Facebook", "preview"),
"meta-externalagent": ("Meta", "preview"),
"skypeuripreview": ("Skype", "preview"),
"bingpreview": ("BingPreview", "preview"),
"pinterest": ("Pinterest", "preview"),
"embedly": ("Embedly", "preview"),
"iframely": ("Iframely", "preview"),
"discordbot": ("Discord", "preview"),
"slackbot": ("Slack", "preview"),
"telegrambot": ("Telegram", "preview"),
"twitterbot": ("Twitter", "preview"),
"linkedinbot": ("LinkedIn", "preview"),
"whatsapp": ("WhatsApp", "preview"),
"ahrefsbot": ("AhrefsBot", "search"),
"mj12bot": ("MJ12bot", "analytics"),
"facebookexternalhit": ("Facebook", "social"),
"meta-externalagent": ("Meta-ExternalAgent", "ai"),
"meta-externalfetcher": ("Meta-ExternalFetcher", "ai"),
"meta-webindexer": ("Meta-WebIndexer", "search"),
"facebookbot": ("FacebookBot", "ai"),
"bingpreview": ("BingPreview", "search"),
"pinterest": ("Pinterest", "social"),
"embedly": ("Embedly", "social"),
"iframely": ("Iframely", "social"),
"discordbot": ("Discord", "social"),
"slackbot": ("Slack", "social"),
"telegrambot": ("Telegram", "social"),
"twitterbot": ("Twitter", "social"),
"linkedinbot": ("LinkedIn", "social"),
"whatsapp": ("WhatsApp", "social"),
"headlesschrome": ("HeadlessChrome", "spider"),
"phantomjs": ("PhantomJS", "spider"),
"uptimerobot": ("UptimeRobot", "spider"),
"pingdom": ("Pingdom", "spider"),
"uptimerobot": ("UptimeRobot", "analytics"),
"pingdom": ("Pingdom", "analytics"),
}
#: Pretty suffixes for the kinds more precise than a generic spider.
KIND_LABEL = {"ai": "AI", "search": "search", "preview": "preview"}
KIND_LABEL = {
"ai": "AI",
"search": "search",
"social": "social",
"analytics": "analytics",
}
#: Providers with more than one crawler product; the kind label is kept
#: only where it distinguishes siblings within the group.
PROVIDERS = [
(
#: Crawler product families: provider -> the display names of its bots.
#: The kind label is kept only where it distinguishes siblings within a family.
PROVIDERS = {
"Google": frozenset({
"Googlebot",
"Google-Extended",
"GoogleOther",
@@ -67,19 +78,33 @@ PROVIDERS = [
"AdsBot-Google",
"APIs-Google",
"Storebot-Google",
"DuplexWeb-Google",
),
("ClaudeBot", "Claude-User", "Claude-SearchBot"),
("GPTBot", "OAI-SearchBot", "ChatGPT-User"),
("PerplexityBot", "Perplexity-User"),
("Amazonbot", "Amzn-SearchBot"),
("Bingbot", "BingPreview"),
]
}),
"Anthropic": frozenset({"ClaudeBot", "Claude-User", "Claude-SearchBot"}),
"OpenAI": frozenset({"GPTBot", "OAI-SearchBot", "ChatGPT-User"}),
"Perplexity": frozenset({"PerplexityBot", "Perplexity-User"}),
"Amazon": frozenset({"Amazonbot", "Amzn-SearchBot"}),
"Microsoft": frozenset({"Bingbot", "BingPreview"}),
"Meta": frozenset({
"Facebook",
"FacebookBot",
"Meta-ExternalAgent",
"Meta-ExternalFetcher",
"Meta-WebIndexer",
}),
}
#: Bot names whose kind label is displayed, computed from the groups.
#: Reverse lookup: bot display name -> provider.
PROVIDER_OF = {
name: provider for provider, names in PROVIDERS.items() for name in names
}
#: Reverse lookup: bot display name -> kind.
NAME_KIND = {name: kind for name, kind in BOTS.values()}
#: Bot names whose kind label is displayed: those in families with mixed kinds.
LABELED = {
name
for group in PROVIDERS
if len({kind for n, kind in BOTS.values() if n in group}) > 1
for name in group
for names in PROVIDERS.values()
if len({NAME_KIND[name] for name in names}) > 1
for name in names
}
+47 -17
View File
@@ -4,18 +4,19 @@ import re
from dataclasses import dataclass
from functools import lru_cache
from .bots import BOTS, KIND_LABEL, LABELED
from .bots import BOTS, KIND_LABEL, LABELED, PROVIDER_OF
from .clients import BROWSERS, SAMSUNG, SAMSUNG_SERIES
@dataclass(frozen=True)
class UA:
pretty: str
engine: str
os: str
bot: str
kind: str
url: str
pretty: str = ""
engine: str = ""
os: str = ""
bot: str = ""
kind: str = ""
url: str = ""
provider: str = ""
#: Fallback for unknown crawlers: a product token whose name says so.
@@ -48,7 +49,7 @@ def url(ua: str) -> str:
def bot(ua: str) -> tuple[str, str]:
"""(display name, kind) of the crawler/previewer the UA claims."""
"""(display name, kind) of the crawler/unfurler the UA claims."""
low = ua.lower()
m = BOTS_RE.search(low)
if m:
@@ -98,6 +99,19 @@ def browser(ua: str) -> str:
#: Engines that differ from the Chromium default for recognized browsers.
ENGINES = {"Firefox": "Gecko", "LibreWolf": "Gecko", "Safari": "Safari"}
#: Oldest plausible major versions (≈2023 releases). These browsers
#: auto-update, so anything older is a scanner's frozen string, not a real
#: installation. Safari is OS-tied and exempt; old Macs genuinely run it.
#: Bump the floors every few years.
ANCIENT = {"Firefox": 108, "Chrome": 108, "Edge": 108, "Opera": 95}
def spoofed(b: str) -> bool:
"""True when a ``Browser/major`` claims an impossibly old version."""
name, _, ver = b.partition("/")
floor = ANCIENT.get(name)
return floor is not None and ver.isdigit() and int(ver) < floor
def engine(b: str) -> str:
"""Engine for a ``Browser/major`` result; Chromium is the modern default."""
@@ -151,20 +165,33 @@ def uaparse(ua: str) -> UA:
would just flush the cache.
"""
if not ua or not ua.strip() or ua in ("-", "null"):
return UA("", "", "", "", "", "")
return UA()
name, kind = bot(ua)
if name:
# The browser/OS in crawler UAs is a disguise; the bot identity is
# the relevant information, so ``engine`` and ``os`` are left empty.
label = KIND_LABEL.get(kind, "") if name in LABELED else ""
pretty = f"{name} ({label})" if label else name
return UA(pretty, "", "", name, kind, url(ua))
return UA(
pretty=pretty, bot=name, kind=kind, url=url(ua),
provider=PROVIDER_OF.get(name, ""),
)
return _parse_client(ua)
@lru_cache(maxsize=1024)
def _parse_client(ua: str) -> UA:
"""Browser/client parsing behind the cache; ``uaparse`` filters bots out."""
r = _client(ua)
# Frozen ancient browser strings are scanners/scripts, not users: show
# the claimed browser, but mark it and drop the fake engine/os/kind.
if r.kind == "browser" and spoofed(browser(ua)):
return UA(pretty=f"{r.pretty} (spoofed)")
return r
def _client(ua: str) -> UA:
"""Browser/client parsing; spoof marking is done by the caller."""
# Non-browser HTTP clients ("python-requests/2.32.5", "curl/8.0",
# "pip/24.3.1 {json…}"): the first product token, plus the OS when
@@ -175,22 +202,22 @@ def _parse_client(ua: str) -> UA:
os_name = os(ua)
if os_name and os_name not in pretty:
pretty = f"{pretty} {os_name}"
return UA(pretty, "", os_name, "", "", "")
return UA(pretty=pretty, os=os_name)
# HarmonyOS carries an "Android" compatibility token, so it must be
# detected before Android.
if "OpenHarmony" in ua or "HarmonyOS" in ua or "ArkWeb" in ua:
b = browser(ua)
return UA(
f"{b} HarmonyOS" if b else "HarmonyOS", "ArkWeb", "HarmonyOS",
"", "browser", "",
pretty=f"{b} HarmonyOS" if b else "HarmonyOS",
engine="ArkWeb", os="HarmonyOS", kind="browser",
)
if "iPhone" in ua or "iPad" in ua:
device = "iPhone" if "iPhone" in ua else "iPad"
m = re.search(r"OS (\d+)", ua)
pretty = f"{device} iOS {m.group(1)}" if m else device
return UA(pretty, "Safari", "iOS", "", "browser", "")
return UA(pretty=pretty, engine="Safari", os="iOS", kind="browser")
m = re.search(r"Android ([\d.]+)", ua)
if m:
@@ -208,14 +235,17 @@ def _parse_client(ua: str) -> UA:
if token and token not in ("wv", "Mobile", "Tablet"):
model = model_name(token)
parts = [p for p in (b, model or f"Android {m.group(1)}") if p]
return UA(" ".join(parts), engine(b), "Android", "", "browser", "")
return UA(
pretty=" ".join(parts), engine=engine(b), os="Android",
kind="browser",
)
b = browser(ua)
os_name = os(ua)
pretty = f"{b} {os_name}".strip()
return UA(pretty or ua, engine(b), os_name, "", "browser", "")
return UA(pretty=pretty or ua, engine=engine(b), os=os_name, kind="browser")
def is_bot(ua: str) -> bool:
"""True when the UA claims a crawler or link-preview identity."""
"""True when the UA claims a crawler or link-unfurling identity."""
return bool(uaparse(ua).bot)