Rework bot taxonomy: provider field, social/analytics kinds, spoof detection
- Distinguish Meta's five documented crawlers (Facebook, FacebookBot, Meta-ExternalAgent/ExternalFetcher/WebIndexer) with distinct kinds - Rename kind preview -> social for link-sharing unfurlers; BingPreview moves to search as it is a search-engine feature - Add analytics kind for monitoring and ad/SEO measurement crawlers (UptimeRobot, Pingdom, MJ12bot, Mediapartners-Google, AdsBot-Google) - Reclassify Google family by main product use: Storebot-Google and Feedfetcher-Google/InspectionTool are search, Read-Aloud and GoogleOther are ai; AhrefsBot is search - Drop retired entries: DuplexWeb-Google, SkypeUriPreview, PhantomJS - Detect the Qualys SSL Labs scanner by its frozen exact UA string - Mark ancient (pre-2023) auto-updating browser claims with " (spoofed)" in pretty, leaving engine/os/kind empty - UA dataclass: all fields default to "", new provider field, constructions use kwargs - PROVIDERS is now dict[provider, frozenset[names]] with derived PROVIDER_OF and NAME_KIND reverse lookups - README and prettytable.py updated to match
This commit is contained in:
@@ -29,8 +29,9 @@ r.url # ""
|
||||
r = uaparse("Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/134.0.6885.65 Mobile Safari/537.36; compatible; facebookexternalhit/1.1; +http://www.facebook.com/externalhit_uatext.php")
|
||||
|
||||
r.pretty # "Facebook"
|
||||
r.kind # "preview"
|
||||
r.kind # "social"
|
||||
r.bot # "Facebook"
|
||||
r.provider # "Meta"
|
||||
r.url # "http://www.facebook.com/externalhit_uatext.php"
|
||||
```
|
||||
|
||||
@@ -38,19 +39,22 @@ r.url # "http://www.facebook.com/externalhit_uatext.php"
|
||||
|
||||
`uaparse(ua)` returns a frozen `UA` dataclass:
|
||||
|
||||
| Field | Content |
|
||||
| -------- | ------------------------------------------------------------------------------------------------ |
|
||||
| `pretty` | Compact display string (below); `""` for empty/missing UAs, the raw UA when unrecognized |
|
||||
| `engine` | `"Chromium"`, `"Gecko"`, `"Safari"`, `"ArkWeb"` (HarmonyOS), or `""` |
|
||||
| `os` | `"Windows"`, `"macOS"`, `"Linux"`, `"iOS"`, `"Android"`, `"HarmonyOS"`, or `""` |
|
||||
| `bot` | Crawler/previewer display name, or `""` |
|
||||
| `kind` | `"browser"`, `"ai"`, `"search"`, `"preview"`, `"spider"`, or `""` (scripts/HTTP libraries) |
|
||||
| `url` | The crawler's info URL (`+https://…` pointer), or `""`; not part of `pretty` — link it in the UI |
|
||||
| Field | Content |
|
||||
| -------- | ------------------------------------------------------------------------------------------------- |
|
||||
| pretty | Compact display string (below); empty for empty/missing UAs, the raw UA when unrecognized |
|
||||
| engine | Chromium, Gecko, Safari, ArkWeb (HarmonyOS), or empty |
|
||||
| os | Windows, macOS, Linux, iOS, Android, HarmonyOS, or empty |
|
||||
| bot | Crawler/unfurler display name, or empty |
|
||||
| kind | browser, ai, search, social, analytics, spider, or empty (scripts/HTTP libraries) |
|
||||
| url | The crawler's info URL (the +https://… pointer), or empty; not part of pretty — link it in the UI |
|
||||
| provider | The bot's provider for known crawler families (Meta, Google, OpenAI, ...), or empty |
|
||||
|
||||
`os` is the major OS only, no version — meant for things like offering OS-specific downloads. `engine` is derived from the browser identity: every recognized browser is Chromium except Firefox/LibreWolf (Gecko) and Safari and all of iOS (Safari's engine is all Apple allows there); HarmonyOS browsers run ArkWeb. Both are left empty for crawlers: the browser and OS in a disguised crawler UA are part of the disguise.
|
||||
|
||||
`kind` is `"browser"` for Mozilla-format UAs with no bot token, `"ai"` for training-data and AI-assistant fetchers (GPTBot, ClaudeBot, Google-Extended, ...), `"search"` for search-engine indexing (Googlebot, Bingbot, ...), `"preview"` for social link-preview fetchers (Facebook, WhatsApp, Slack, ...), `"spider"` for generic or unknown crawlers, and `""` for scripts and HTTP libraries.
|
||||
|
||||
The kind field describes our detection of visitor type: browser for actual browsers, ai for AI training collectors, agents and user-initiated fetches (GPTBot, ClaudeBot, ChatGPT-User, Google-Extended, ...), search for search-engine indexing (Googlebot, Bingbot, ...), social for link-sharing unfurlers (Facebook, WhatsApp, Slack, ...), analytics for monitoring and site-analytics crawlers (UptimeRobot, AdsBot-Google, MJ12bot), spider for generic or unknown crawlers, and empty for scripts and HTTP libraries. Any value other than browser means the visitor is automated.
|
||||
|
||||
`pretty` is intended to be shown directly:
|
||||
|
||||
- Desktop: `Chrome/152 Windows`, `Safari/18 macOS`
|
||||
@@ -90,11 +94,11 @@ The table below compares representative results. uarite shows `r.pretty`; the ua
|
||||
| Huawei HarmonyOS phone | HuaweiBrowser/6 HarmonyOS | Huawei Browser/6 Android❌ Huawei Browser |
|
||||
| GPTBot | GPTBot (AI) | GPTBot/1 Spider |
|
||||
| Googlebot (disguised) | Googlebot (search) | Googlebot/2 Android❌ Spider |
|
||||
| Facebook preview (disguised) | Facebook | FacebookBot/1 Android Pixel 7 ❌ |
|
||||
| Meta crawler (disguised) | Meta | Chrome/145 Windows ❌ |
|
||||
| Facebook preview (disguised) | Facebook (social) | FacebookBot/1 Android Pixel 7 ❌ |
|
||||
| Meta crawler (disguised) | Meta-ExternalAgent (AI) | Chrome/145 Windows ❌ |
|
||||
| WhatsApp preview | WhatsApp | WhatsApp/10 Spider |
|
||||
| Bytespider | Bytespider | Bytespider/ Android❌ Generic Smartphone |
|
||||
| BingPreview | BingPreview (preview) | BingPreview/1 Windows❌ Spider |
|
||||
| BingPreview | BingPreview | BingPreview/1 Windows❌ Spider |
|
||||
| AhrefsBot | AhrefsBot | AhrefsBot/7 Spider |
|
||||
| python-requests | python-requests/2.32.5 | Python Requests/2 |
|
||||
|
||||
|
||||
@@ -55,10 +55,6 @@ CASES = [
|
||||
"Googlebot (disguised)",
|
||||
"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.84 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
|
||||
),
|
||||
(
|
||||
"Claude-SearchBot (disguised)",
|
||||
"Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/148.0.3245.171 Mobile Safari/537.36; compatible; Claude-SearchBot/1.0; +https://www.anthropic.com/claude-searchbot",
|
||||
),
|
||||
(
|
||||
"Facebook preview (disguised)",
|
||||
"Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/134.0.6885.65 Mobile Safari/537.36; compatible; facebookexternalhit/1.1; +http://www.facebook.com/externalhit_uatext.php",
|
||||
|
||||
+67
-42
@@ -1,19 +1,24 @@
|
||||
"""Crawler and link-preview tables."""
|
||||
"""Crawler and link-unfurler tables."""
|
||||
|
||||
#: Lowercase UA substring to (display name, kind). Ordered: first match
|
||||
#: wins, so overlapping names go from most to least specific.
|
||||
BOTS = {
|
||||
# Qualys SSL Labs scanner: frozen on this exact Firefox/45 string since
|
||||
# ~2016; the pinned Gecko date makes the substring distinctive.
|
||||
"mozilla/5.0 (x11; linux x86_64; rv:45.0) gecko/20100101 firefox/45.0": (
|
||||
"Qualys SSL Labs",
|
||||
"spider",
|
||||
),
|
||||
"feedfetcher-google": ("Feedfetcher-Google", "search"),
|
||||
"google-inspectiontool": ("Google-InspectionTool", "search"),
|
||||
"google-read-aloud": ("Google-Read-Aloud", "search"),
|
||||
"mediapartners-google": ("Mediapartners-Google", "spider"),
|
||||
"adsbot-google": ("AdsBot-Google", "spider"),
|
||||
"google-read-aloud": ("Google-Read-Aloud", "ai"),
|
||||
"mediapartners-google": ("Mediapartners-Google", "analytics"),
|
||||
"adsbot-google": ("AdsBot-Google", "analytics"),
|
||||
"apis-google": ("APIs-Google", "spider"),
|
||||
"storebot-google": ("Storebot-Google", "spider"),
|
||||
"duplexweb-google": ("DuplexWeb-Google", "spider"),
|
||||
"storebot-google": ("Storebot-Google", "search"),
|
||||
"google-extended": ("Google-Extended", "ai"),
|
||||
"googlebot": ("Googlebot", "search"),
|
||||
"googleother": ("GoogleOther", "spider"),
|
||||
"googleother": ("GoogleOther", "ai"),
|
||||
"bingbot": ("Bingbot", "search"),
|
||||
"applebot": ("Applebot", "search"),
|
||||
"gptbot": ("GPTBot", "ai"),
|
||||
@@ -29,34 +34,40 @@ BOTS = {
|
||||
"reflectionbot": ("Reflectionbot", "ai"),
|
||||
"amzn-searchbot": ("Amzn-SearchBot", "search"),
|
||||
"amazonbot": ("Amazonbot", "search"),
|
||||
"ahrefsbot": ("AhrefsBot", "spider"),
|
||||
"mj12bot": ("MJ12bot", "spider"),
|
||||
"facebookexternalhit": ("Facebook", "preview"),
|
||||
"meta-externalagent": ("Meta", "preview"),
|
||||
"skypeuripreview": ("Skype", "preview"),
|
||||
"bingpreview": ("BingPreview", "preview"),
|
||||
"pinterest": ("Pinterest", "preview"),
|
||||
"embedly": ("Embedly", "preview"),
|
||||
"iframely": ("Iframely", "preview"),
|
||||
"discordbot": ("Discord", "preview"),
|
||||
"slackbot": ("Slack", "preview"),
|
||||
"telegrambot": ("Telegram", "preview"),
|
||||
"twitterbot": ("Twitter", "preview"),
|
||||
"linkedinbot": ("LinkedIn", "preview"),
|
||||
"whatsapp": ("WhatsApp", "preview"),
|
||||
"ahrefsbot": ("AhrefsBot", "search"),
|
||||
"mj12bot": ("MJ12bot", "analytics"),
|
||||
"facebookexternalhit": ("Facebook", "social"),
|
||||
"meta-externalagent": ("Meta-ExternalAgent", "ai"),
|
||||
"meta-externalfetcher": ("Meta-ExternalFetcher", "ai"),
|
||||
"meta-webindexer": ("Meta-WebIndexer", "search"),
|
||||
"facebookbot": ("FacebookBot", "ai"),
|
||||
"bingpreview": ("BingPreview", "search"),
|
||||
"pinterest": ("Pinterest", "social"),
|
||||
"embedly": ("Embedly", "social"),
|
||||
"iframely": ("Iframely", "social"),
|
||||
"discordbot": ("Discord", "social"),
|
||||
"slackbot": ("Slack", "social"),
|
||||
"telegrambot": ("Telegram", "social"),
|
||||
"twitterbot": ("Twitter", "social"),
|
||||
"linkedinbot": ("LinkedIn", "social"),
|
||||
"whatsapp": ("WhatsApp", "social"),
|
||||
"headlesschrome": ("HeadlessChrome", "spider"),
|
||||
"phantomjs": ("PhantomJS", "spider"),
|
||||
"uptimerobot": ("UptimeRobot", "spider"),
|
||||
"pingdom": ("Pingdom", "spider"),
|
||||
"uptimerobot": ("UptimeRobot", "analytics"),
|
||||
"pingdom": ("Pingdom", "analytics"),
|
||||
}
|
||||
|
||||
#: Pretty suffixes for the kinds more precise than a generic spider.
|
||||
KIND_LABEL = {"ai": "AI", "search": "search", "preview": "preview"}
|
||||
KIND_LABEL = {
|
||||
"ai": "AI",
|
||||
"search": "search",
|
||||
"social": "social",
|
||||
"analytics": "analytics",
|
||||
}
|
||||
|
||||
#: Providers with more than one crawler product; the kind label is kept
|
||||
#: only where it distinguishes siblings within the group.
|
||||
PROVIDERS = [
|
||||
(
|
||||
#: Crawler product families: provider -> the display names of its bots.
|
||||
#: The kind label is kept only where it distinguishes siblings within a family.
|
||||
PROVIDERS = {
|
||||
"Google": frozenset({
|
||||
"Googlebot",
|
||||
"Google-Extended",
|
||||
"GoogleOther",
|
||||
@@ -67,19 +78,33 @@ PROVIDERS = [
|
||||
"AdsBot-Google",
|
||||
"APIs-Google",
|
||||
"Storebot-Google",
|
||||
"DuplexWeb-Google",
|
||||
),
|
||||
("ClaudeBot", "Claude-User", "Claude-SearchBot"),
|
||||
("GPTBot", "OAI-SearchBot", "ChatGPT-User"),
|
||||
("PerplexityBot", "Perplexity-User"),
|
||||
("Amazonbot", "Amzn-SearchBot"),
|
||||
("Bingbot", "BingPreview"),
|
||||
]
|
||||
}),
|
||||
"Anthropic": frozenset({"ClaudeBot", "Claude-User", "Claude-SearchBot"}),
|
||||
"OpenAI": frozenset({"GPTBot", "OAI-SearchBot", "ChatGPT-User"}),
|
||||
"Perplexity": frozenset({"PerplexityBot", "Perplexity-User"}),
|
||||
"Amazon": frozenset({"Amazonbot", "Amzn-SearchBot"}),
|
||||
"Microsoft": frozenset({"Bingbot", "BingPreview"}),
|
||||
"Meta": frozenset({
|
||||
"Facebook",
|
||||
"FacebookBot",
|
||||
"Meta-ExternalAgent",
|
||||
"Meta-ExternalFetcher",
|
||||
"Meta-WebIndexer",
|
||||
}),
|
||||
}
|
||||
|
||||
#: Bot names whose kind label is displayed, computed from the groups.
|
||||
#: Reverse lookup: bot display name -> provider.
|
||||
PROVIDER_OF = {
|
||||
name: provider for provider, names in PROVIDERS.items() for name in names
|
||||
}
|
||||
|
||||
#: Reverse lookup: bot display name -> kind.
|
||||
NAME_KIND = {name: kind for name, kind in BOTS.values()}
|
||||
|
||||
#: Bot names whose kind label is displayed: those in families with mixed kinds.
|
||||
LABELED = {
|
||||
name
|
||||
for group in PROVIDERS
|
||||
if len({kind for n, kind in BOTS.values() if n in group}) > 1
|
||||
for name in group
|
||||
for names in PROVIDERS.values()
|
||||
if len({NAME_KIND[name] for name in names}) > 1
|
||||
for name in names
|
||||
}
|
||||
|
||||
+47
-17
@@ -4,18 +4,19 @@ import re
|
||||
from dataclasses import dataclass
|
||||
from functools import lru_cache
|
||||
|
||||
from .bots import BOTS, KIND_LABEL, LABELED
|
||||
from .bots import BOTS, KIND_LABEL, LABELED, PROVIDER_OF
|
||||
from .clients import BROWSERS, SAMSUNG, SAMSUNG_SERIES
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class UA:
|
||||
pretty: str
|
||||
engine: str
|
||||
os: str
|
||||
bot: str
|
||||
kind: str
|
||||
url: str
|
||||
pretty: str = ""
|
||||
engine: str = ""
|
||||
os: str = ""
|
||||
bot: str = ""
|
||||
kind: str = ""
|
||||
url: str = ""
|
||||
provider: str = ""
|
||||
|
||||
|
||||
#: Fallback for unknown crawlers: a product token whose name says so.
|
||||
@@ -48,7 +49,7 @@ def url(ua: str) -> str:
|
||||
|
||||
|
||||
def bot(ua: str) -> tuple[str, str]:
|
||||
"""(display name, kind) of the crawler/previewer the UA claims."""
|
||||
"""(display name, kind) of the crawler/unfurler the UA claims."""
|
||||
low = ua.lower()
|
||||
m = BOTS_RE.search(low)
|
||||
if m:
|
||||
@@ -98,6 +99,19 @@ def browser(ua: str) -> str:
|
||||
#: Engines that differ from the Chromium default for recognized browsers.
|
||||
ENGINES = {"Firefox": "Gecko", "LibreWolf": "Gecko", "Safari": "Safari"}
|
||||
|
||||
#: Oldest plausible major versions (≈2023 releases). These browsers
|
||||
#: auto-update, so anything older is a scanner's frozen string, not a real
|
||||
#: installation. Safari is OS-tied and exempt; old Macs genuinely run it.
|
||||
#: Bump the floors every few years.
|
||||
ANCIENT = {"Firefox": 108, "Chrome": 108, "Edge": 108, "Opera": 95}
|
||||
|
||||
|
||||
def spoofed(b: str) -> bool:
|
||||
"""True when a ``Browser/major`` claims an impossibly old version."""
|
||||
name, _, ver = b.partition("/")
|
||||
floor = ANCIENT.get(name)
|
||||
return floor is not None and ver.isdigit() and int(ver) < floor
|
||||
|
||||
|
||||
def engine(b: str) -> str:
|
||||
"""Engine for a ``Browser/major`` result; Chromium is the modern default."""
|
||||
@@ -151,20 +165,33 @@ def uaparse(ua: str) -> UA:
|
||||
would just flush the cache.
|
||||
"""
|
||||
if not ua or not ua.strip() or ua in ("-", "null"):
|
||||
return UA("", "", "", "", "", "")
|
||||
return UA()
|
||||
name, kind = bot(ua)
|
||||
if name:
|
||||
# The browser/OS in crawler UAs is a disguise; the bot identity is
|
||||
# the relevant information, so ``engine`` and ``os`` are left empty.
|
||||
label = KIND_LABEL.get(kind, "") if name in LABELED else ""
|
||||
pretty = f"{name} ({label})" if label else name
|
||||
return UA(pretty, "", "", name, kind, url(ua))
|
||||
return UA(
|
||||
pretty=pretty, bot=name, kind=kind, url=url(ua),
|
||||
provider=PROVIDER_OF.get(name, ""),
|
||||
)
|
||||
return _parse_client(ua)
|
||||
|
||||
|
||||
@lru_cache(maxsize=1024)
|
||||
def _parse_client(ua: str) -> UA:
|
||||
"""Browser/client parsing behind the cache; ``uaparse`` filters bots out."""
|
||||
r = _client(ua)
|
||||
# Frozen ancient browser strings are scanners/scripts, not users: show
|
||||
# the claimed browser, but mark it and drop the fake engine/os/kind.
|
||||
if r.kind == "browser" and spoofed(browser(ua)):
|
||||
return UA(pretty=f"{r.pretty} (spoofed)")
|
||||
return r
|
||||
|
||||
|
||||
def _client(ua: str) -> UA:
|
||||
"""Browser/client parsing; spoof marking is done by the caller."""
|
||||
|
||||
# Non-browser HTTP clients ("python-requests/2.32.5", "curl/8.0",
|
||||
# "pip/24.3.1 {json…}"): the first product token, plus the OS when
|
||||
@@ -175,22 +202,22 @@ def _parse_client(ua: str) -> UA:
|
||||
os_name = os(ua)
|
||||
if os_name and os_name not in pretty:
|
||||
pretty = f"{pretty} {os_name}"
|
||||
return UA(pretty, "", os_name, "", "", "")
|
||||
return UA(pretty=pretty, os=os_name)
|
||||
|
||||
# HarmonyOS carries an "Android" compatibility token, so it must be
|
||||
# detected before Android.
|
||||
if "OpenHarmony" in ua or "HarmonyOS" in ua or "ArkWeb" in ua:
|
||||
b = browser(ua)
|
||||
return UA(
|
||||
f"{b} HarmonyOS" if b else "HarmonyOS", "ArkWeb", "HarmonyOS",
|
||||
"", "browser", "",
|
||||
pretty=f"{b} HarmonyOS" if b else "HarmonyOS",
|
||||
engine="ArkWeb", os="HarmonyOS", kind="browser",
|
||||
)
|
||||
|
||||
if "iPhone" in ua or "iPad" in ua:
|
||||
device = "iPhone" if "iPhone" in ua else "iPad"
|
||||
m = re.search(r"OS (\d+)", ua)
|
||||
pretty = f"{device} iOS {m.group(1)}" if m else device
|
||||
return UA(pretty, "Safari", "iOS", "", "browser", "")
|
||||
return UA(pretty=pretty, engine="Safari", os="iOS", kind="browser")
|
||||
|
||||
m = re.search(r"Android ([\d.]+)", ua)
|
||||
if m:
|
||||
@@ -208,14 +235,17 @@ def _parse_client(ua: str) -> UA:
|
||||
if token and token not in ("wv", "Mobile", "Tablet"):
|
||||
model = model_name(token)
|
||||
parts = [p for p in (b, model or f"Android {m.group(1)}") if p]
|
||||
return UA(" ".join(parts), engine(b), "Android", "", "browser", "")
|
||||
return UA(
|
||||
pretty=" ".join(parts), engine=engine(b), os="Android",
|
||||
kind="browser",
|
||||
)
|
||||
|
||||
b = browser(ua)
|
||||
os_name = os(ua)
|
||||
pretty = f"{b} {os_name}".strip()
|
||||
return UA(pretty or ua, engine(b), os_name, "", "browser", "")
|
||||
return UA(pretty=pretty or ua, engine=engine(b), os=os_name, kind="browser")
|
||||
|
||||
|
||||
def is_bot(ua: str) -> bool:
|
||||
"""True when the UA claims a crawler or link-preview identity."""
|
||||
"""True when the UA claims a crawler or link-unfurling identity."""
|
||||
return bool(uaparse(ua).bot)
|
||||
|
||||
Reference in New Issue
Block a user