13 Commits
Author SHA1 Message Date
LeoVasanko a225bacfc4 Add uarite-js: JavaScript/TypeScript port published to npm. Bump version. 2026-09-09 23:35:59 +00:00
LeoVasanko f25bfafbc7 README 2026-09-09 19:39:33 +00:00
LeoVasanko 7ae4fab180 README 2026-09-09 19:04:29 +00:00
LeoVasanko 5bbd9a5484 Updated docs with fastuaparser and new benchmarks, cleaner structure. 2026-09-09 18:59:02 +00:00
LeoVasanko 319aa52a6f README, and support script cleanup. 2026-09-08 23:47:50 +00:00
LeoVasanko a7fdb4ddf3 Rework benchmark: cold-cache 50/50 speed, ua-parser native backends, SVG throughput plot 2026-09-08 23:36:31 +00:00
LeoVasanko 887d492598 Normalize bot pretty names; drop undocumented FacebookBot
PRETTY_OVERRIDE in bots.py replaces name-plus-label composition with full per-bot pretty strings. The messy X-Google / Google-X names print as "Google X (kind)"; Googlebot keeps its canonical name. FacebookBot is no longer in Meta's crawler documentation, so it falls back to generic spider detection; Facebook rejoins provider Meta but prints bare, like the other social unfurlers.
2026-09-08 23:36:31 +00:00
LeoVasanko dd59fcb832 Accept None in uaparse's type signature (already handled) 2026-09-08 23:29:05 +00:00
LeoVasanko 0bdbe9aaf9 Cache uaparse fully; drop the bot field and is_bot 2026-09-08 23:23:02 +00:00
LeoVasanko 84b9e27ced Rework bot taxonomy: provider field, social/analytics kinds, spoof detection
- Distinguish Meta's five documented crawlers (Facebook, FacebookBot,
  Meta-ExternalAgent/ExternalFetcher/WebIndexer) with distinct kinds
- Rename kind preview -> social for link-sharing unfurlers; BingPreview
  moves to search as it is a search-engine feature
- Add analytics kind for monitoring and ad/SEO measurement crawlers
  (UptimeRobot, Pingdom, MJ12bot, Mediapartners-Google, AdsBot-Google)
- Reclassify Google family by main product use: Storebot-Google and
  Feedfetcher-Google/InspectionTool are search, Read-Aloud and
  GoogleOther are ai; AhrefsBot is search
- Drop retired entries: DuplexWeb-Google, SkypeUriPreview, PhantomJS
- Detect the Qualys SSL Labs scanner by its frozen exact UA string
- Mark ancient (pre-2023) auto-updating browser claims with
  " (spoofed)" in pretty, leaving engine/os/kind empty
- UA dataclass: all fields default to "", new provider field,
  constructions use kwargs
- PROVIDERS is now dict[provider, frozenset[names]] with derived
  PROVIDER_OF and NAME_KIND reverse lookups
- README and prettytable.py updated to match
2026-09-08 23:20:06 +00:00
LeoVasanko b70c8f313c Additional Samsung model detection. 2026-09-05 06:00:30 +00:00
LeoVasanko 4d6f1a2afe Updated README. 2026-09-05 05:12:22 +00:00
LeoVasanko 99dea2b9eb Add engine and os fields to UA, restructure README prose, restyle pyproject 2026-09-05 04:49:07 +00:00
20 changed files with 2063 additions and 233 deletions
+8 -1
View File
@@ -2,10 +2,17 @@
__pycache__/
*.py[oc]
build/
dist/
/dist/
wheels/
*.egg-info
.*
!.gitignore
!*/.prettierrc.json
*.lock
# JS build artifacts
node_modules/
package-lock.json
uarite-js/dist/*
!uarite-js/dist/uarite.min.js
+67 -76
View File
@@ -1,119 +1,110 @@
# uarite
# User-Agent Parsing Done Right
User-Agent parsing in Python has a long lineage. ua-parser is the official Python implementation of the ua-parser project, built around uap-core: the regex database extracted from BrowserScope's original parser and shared by implementations in many languages. user-agents wraps ua-parser with higher-level device and capability detection; its last release was in 2020. user-agent-parser is a separate implementation first released in 2022 and substantially updated in 2026, taking its own approach rather than building on uap-core. None of the three has further dependencies, but the regex databases weigh something: ua-parser installs at about 499 KB (531 KB with user-agents on top), user-agent-parser at 166 KB.
Fast and accurate handling of modern browsers and crawlers. Despite its light weight, uarite identifies both browsers and crawlers more accurately than any competing implementation tested here. Despite being pure Python, it outperforms ua-parser's C++/Rust variants. We also provide a [JavaScript uarite](https://www.npmjs.com/package/@vasanko/uarite) with exact same output.
It returns structured classification, but also the thing most applications eventually need: **a short pretty description**.
## Usage
Add it to your project:
```sh
uv add uarite
```
This module is another take on the same problem: a small, dependency-free, compact pure-Python parser — 29 KB installed. It returns structured classifications, but also the thing most applications eventually need: a short human-readable description.
It is particularly aimed at server-side analytics, where correctly recognizing crawlers and modern reduced User-Agents matters. It detects disguised crawlers, distinguishes AI/search/preview traffic, handles HarmonyOS and bots without calling them Android, resolves common device model codes and falls back to reasonable output even when all else fails.
## Usage
```python
from uarite import uaparse
r = uaparse("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36")
r = uaparse("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36")
r.pretty # "Chrome/152 Windows"
r.engine # "Chromium"
r.os # "Windows"
r.kind # "browser"
r.bot # ""
r.url # ""
r = uaparse("Mozilla/5.0 (Linux; Android 13; Pixel 7) ... Chrome/134.0.6885.65 "
"Mobile Safari/537.36; compatible; facebookexternalhit/1.1; "
"+http://www.facebook.com/externalhit_uatext.php")
r = uaparse("Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/134.0.6885.65 Mobile Safari/537.36; compatible; facebookexternalhit/1.1; +http://www.facebook.com/externalhit_uatext.php")
r.pretty # "Facebook"
r.kind # "preview"
r.bot # "Facebook"
r.kind # "social"
r.provider # "Meta"
r.url # "http://www.facebook.com/externalhit_uatext.php"
```
## Output
`uaparse(ua)` returns a frozen `UA` dataclass:
`uaparse(ua)` returns a `UA` dataclass with string fields. Any field may be empty when the information is unavailable.
| Field | Content |
| -------- | ------------------------------------------------------------------------------------------------ |
| `pretty` | Compact display string (below); `""` for empty/missing UAs, the raw UA when unrecognized |
| `bot` | Crawler/previewer display name, or `""` |
| `kind` | `"browser"`, `"ai"`, `"search"`, `"preview"`, `"spider"`, or `""` (scripts/HTTP libraries) |
| `url` | The crawler's info URL (`+https://…` pointer), or `""`; not part of `pretty` — link it in the UI |
| Field | Content |
| -------- | ------------------------------------------------ |
| pretty | Compact display string; raw UA when unrecognized |
| engine | Chromium, Gecko, Safari, ArkWeb |
| os | Windows, macOS, Linux, iOS, Android, HarmonyOS |
| kind | browser, ai, search, social, analytics, spider |
| url | Crawler information URL |
| provider | Provider of a known crawler family |
`kind` is `"browser"` for Mozilla-format UAs with no bot token, `"ai"` for training-data and AI-assistant fetchers (GPTBot, ClaudeBot, Google-Extended, ...), `"search"` for search-engine indexing (Googlebot, Bingbot, ...), `"preview"` for social link-preview fetchers (Facebook, WhatsApp, Slack, ...), `"spider"` for generic or unknown crawlers, and `""` for scripts and HTTP libraries.
The pretty field is intended for UIs and logs. The url can be attached to it as a link when available.
`pretty` is intended to be shown directly:
The engine and os fields are intentionally broad. The kind field distinguishes browsers from AI collectors, search engines, social previews, monitoring tools, generic spiders, and ordinary HTTP clients. Any non-browser kind represents automated traffic.
- Desktop: `Chrome/152 Windows`, `Safari/18 macOS`
- iPhone/iPad: `iPhone iOS 17` — the device and iOS version, not Safari (the only browser iOS has)
- Android: `Chrome/118 Pixel 6`, or `Chrome/152 Android` when the device is unknown
- Crawlers: `GPTBot (AI)`, `Googlebot (search)`, `Facebook` — the kind suffix appears only where a provider runs crawlers of more than one kind; single-kind providers stay plain
- Scripts: `python-requests/2.32.5`, `pip/24.3.1 Linux`
Detection is necessarily limited by what the User-Agent reveals. Crawlers can masquerade as ordinary browsers or other crawlers, so sites that need stronger identification should use additional methods rather than relying on UA detection alone.
Chrome's reduced Android UA reports the frozen values `Android 10; K`; neither is real device information, so uarite deliberately reports simply `Android`. HarmonyOS compatibility strings are similarly recognized before their misleading Android tokens.
## Comparison
## Performance
There are plenty of UA parsers for Python. This comparison includes the ua-parser family — the official port of the large upstream project, benchmarked with all three of its regex backends (Python, RE2, Rust) — plus user-agents, which builds on it with higher-level device detection, and fastuaparser, a single-file heuristic parser in the same weight class as uarite. Among those excluded: user-agent-parser crashes on real-world strings; httpagentparser misidentifies most bots and browser OSes; device-detector has a massive 26 MB install and extremely slow parsing; and crawlerdetect only distinguishes bots from non-bots. The user-agents package has been unmaintained since 2020 but is kept as a reference point.
All compared parsers cache repeated User-Agents, making cache hits effectively free. The useful difference is therefore the first parse of a new string.
Installed size tracks the design. The minimal heuristic parsers are tiny: fastuaparser installs in 15 kB, uarite in 30 kB. Everything built on ua-parser's regex database starts around half a megabyte, and its accelerated backends add several megabytes of native code.
In our benchmarks, uncached uarite parses take roughly **37 µs**. user-agent-parser is in the same general range at **~7 µs**, while the pure-Python ua-parser/user-agents path takes roughly **130250 µs**.
The benchmark and test scripts are available in the repository's scripts folder.
The cache strategies differ in ways that matter under adversarial traffic. uarite caches only browser/client results (1024-entry LRU): crawlers tend to be unique and would otherwise evict the repeating UAs where caching is useful. user-agent-parser's 512-entry LRU lets a bot storm evict browsers, and user-agents' 200-entry dict clears entirely when full.
### Accuracy
## Accuracy
The table below compares representative User-Agent formats with fastuaparser and ua-parser. The user-agents wrapper produces nearly identical results to ua-parser and is left out.
The main difference is not how many fields can be returned, but what the parser believes the User-Agent actually says.
| Case | uarite¹ | fastuaparser² | ua-parser³ |
| ---------------------------- | ------------------------- | ------------------------- | --------------------------------------------- |
| Chrome, Windows | Chrome/152 Windows | Chrome - Windows | Chrome/152 Windows |
| Safari, macOS | Safari/18 macOS | Safari - Mac OS X | Safari/18 Mac OS X Mac |
| Opera, Linux | Opera/106 Linux | Opera - Linux | Opera/106 Linux |
| Chrome, Android (no model) | Chrome/152 Android | Chrome - Android Mobile | Chrome Mobile/152 Android K❌ |
| Chrome, Android (Pixel) | Chrome/118 Pixel 6 | Chrome - Android Mobile | Chrome Mobile/118 Android Pixel 6 |
| Edge, Android (model code) | Edge/110 Galaxy S7 | Edge - Android Mobile | Edge Mobile/110 Android Samsung SM-G930P |
| Firefox, Android | Firefox/154 Android 15 | Firefox - Android Mobile | Firefox Mobile/154 Android Generic Smartphone |
| Safari, iPhone | iPhone iOS 17 | Safari - iOS Mobile | Mobile Safari/17 iOS iPhone |
| Huawei HarmonyOS phone | HuaweiBrowser/6 HarmonyOS | Chrome - Android Mobile❌ | Huawei Browser/6 Android❌ Huawei Browser |
| GPTBot | GPTBot (AI) | Bot | GPTBot/1 Spider |
| Googlebot (disguised) | Googlebot (search) | Bot | Googlebot/2 Android❌ Spider |
| Facebook preview (disguised) | Facebook | Chrome - Android Mobile❌ | FacebookBot/1 Android Pixel 7 ❌ |
| Meta crawler (disguised) | Meta-ExternalAgent (AI) | Bot | Chrome/145 Windows ❌ |
| WhatsApp preview | WhatsApp | Browser - Other❌ | WhatsApp/10 Spider |
| Bytespider | Bytespider | Bot | Bytespider/ Android❌ Generic Smartphone |
| BingPreview | BingPreview | Browser - Windows❌ | BingPreview/1 Windows❌ Spider |
| AhrefsBot | AhrefsBot | Bot | AhrefsBot/7 Spider |
| python-requests | python-requests/2.32.5 | Other | Python Requests/2 |
For example, reduced Chrome does not really tell us that the device is named `K` or that it runs Android 10; an Android compatibility token does not make HarmonyOS Android; and a Facebook or Google crawler containing a plausible Chrome UA is still a crawler, not a Chrome visitor.
- ❌ marks incorrect data such as an OS from disguise, Android 10; K (compat) on modern devices, or a missed identity
- ¹ `r.pretty` shown as is
- ² `parse_ua(ua)` shown as is
- ³ `{r.user_agent.family}/{r.user_agent.major} {r.os.family} {r.device.family}`
The table below compares representative results. uarite shows `r.pretty`; the ua-parser display strings are assembled from its structured output for comparison. user-agents is omitted: it shares the ua-parser backend and returns virtually identical data in a slightly different structure.
On our modern-browser test set, scoring only fully correct results (family, version, and OS all right): **uarite 100%**, ua-parser and user-agents 79%. The fastuaparser heuristic resolves the family and OS in 91% of cases, but reports no version numbers at all, so it cannot be scored on the full criterion.
| Case | uarite¹ | ua-parser² |
| ---------------------------- | ------------------------- | --------------------------------------------- |
| Chrome, Windows | Chrome/152 Windows | Chrome/152 Windows |
| Safari, macOS | Safari/18 macOS | Safari/18 Mac OS X Mac |
| Opera, Linux | Opera/106 Linux | Opera/106 Linux |
| Chrome, Android (no model) | Chrome/152 Android | Chrome Mobile/152 Android K❌ |
| Chrome, Android (Pixel) | Chrome/118 Pixel 6 | Chrome Mobile/118 Android Pixel 6 |
| Edge, Android (model code) | Edge/110 Galaxy S7 | Edge Mobile/110 Android Samsung SM-G930P |
| Firefox, Android | Firefox/154 Android 15 | Firefox Mobile/154 Android Generic Smartphone |
| Safari, iPhone | iPhone iOS 17 | Mobile Safari/17 iOS iPhone |
| Huawei HarmonyOS phone | HuaweiBrowser/6 HarmonyOS | Huawei Browser/6 Android❌ Huawei Browser |
| GPTBot | GPTBot (AI) | GPTBot/1 Spider |
| Googlebot (disguised) | Googlebot (search) | Googlebot/2 Android❌ Spider |
| Facebook preview (disguised) | Facebook | FacebookBot/1 Android Pixel 7 ❌ |
| Meta crawler (disguised) | Meta | Chrome/145 Windows ❌ |
| WhatsApp preview | WhatsApp | WhatsApp/10 Spider |
| Bytespider | Bytespider | Bytespider/ Android❌ Generic Smartphone |
| BingPreview | BingPreview (preview) | BingPreview/1 Windows❌ Spider |
| AhrefsBot | AhrefsBot | AhrefsBot/7 Spider |
| python-requests | python-requests/2.32.5 | Python Requests/2 |
Crawler detection was tested against [real world UAs](https://github.com/monperrus/crawler-user-agents): **uarite 97%**, fastuaparser 84%, ua-parser 64%, user-agents 60%.
❌ marks an incorrect browser, OS, or device interpretation.
¹ `r.pretty` shown as is
² `{user_agent.family}/{user_agent.major} {os.family} {device.family}`
### Performance
Measured on 100 current browser UAs, family / version / OS accuracy is 80% / 91% / 99% for both ua-parser and user-agents, and 90% / 90% / 100% for uarite — its nominal "misses" are the iPhone rows, where it reports `iPhone iOS 17` rather than Mobile Safari, a deliberate choice since Safari is the only browser iOS has. user-agent-parser is absent from the table: it detected under a third of the crawlers in the test below and crashed on five inputs, so a side-by-side formatting comparison adds little.
Import and first parse takes about **10 ms** for uarite and effectively nothing for fastuaparser, compared with 5090 ms for ua-parser depending on backend.
Crawler detection was also tested against 2163 real-world crawler UAs from [monperrus/crawler-user-agents](https://github.com/monperrus/crawler-user-agents):
![uarite 56 thousand, ua-parser native code variants Rust 26 thousand, RE2 14 thousand, and finally plain Python ua-parser and user-agents 3 thousand](https://git.zi.fi/LeoVasanko/uarite/raw/branch/main/docs/bench-speed.svg)
_User-Agents parsed per second per CPU core, first parse of unseen strings, with equal shares of browser and crawler UAs. One-off setup costs excluded. Cached results and fastuaparser (1 million) are left out of the graph._
| Parser | Detected |
| ----------------- | ------------------------ |
| ua-parser | 63.8% |
| user-agents | 60.1% |
| user-agent-parser | 31.9% (crashed on 5 UAs) |
| uarite | 79.9% |
On raw speed fastuaparser wins: a few string searches per UA, no cache needed. The trade-off shows in the accuracy table above. All other parsers cache results. With cache hits, **uarite reaches about 36 million lookups per second**, compared with about 5 million for ua-parser and 600 000 for user-agents.
The remaining uarite misses are mostly ancient tokenless crawler names and ordinary HTTP libraries, which are intentionally classified as scripts rather than bots.
## Why yet another UA parser
## Design
Rather than relying on a large historical regex database, the parser focuses on modern UA formats and parses them directly, choosing the most specific interpretation available. This keeps the implementation small while handling today's browsers and crawler traffic well. Until now, I had been using those other modules and building my own pretty-UA formatting on top of them, fixing by post processing issues the upstream didn't care of.
uarite uses a small hand-maintained regex/rule database rather than the much larger uap-core dataset. Rules are kept simple enough to audit directly and ordered so that specific identities such as crawlers or HarmonyOS win over browser compatibility tokens. The generic tells are few: a product token containing bot/spider/crawl/scan/verify/check, or an info URL in the UA — real browsers never carry one.
Eventually it became easier to start over with a parser designed around modern traffic. The result is uarite.
Unknown UAs remain visible: `pretty` falls back to the original string rather than discarding them, and malformed input never raises.
Some information simply is not present in a User-Agent. Modern Brave is normally indistinguishable from Chrome without browser-side detection, and iPhone UAs do not contain the device model. uarite prefers an incomplete answer to an invented one.
Hopefully it helps you too. Star my [GitHub](https://github.com/leovasanko/uarite) if it did.
+694
View File
@@ -0,0 +1,694 @@
<?xml version="1.0" encoding="utf-8" standalone="no"?>
<!DOCTYPE svg PUBLIC "-//W3C//DTD SVG 1.1//EN"
"http://www.w3.org/Graphics/SVG/1.1/DTD/svg11.dtd">
<svg xmlns:xlink="http://www.w3.org/1999/xlink" width="758.3677pt" height="86.04pt" viewBox="0 0 758.3677 86.04" xmlns="http://www.w3.org/2000/svg" version="1.1">
<metadata>
<rdf:RDF xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:cc="http://creativecommons.org/ns#" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#">
<cc:Work>
<dc:type rdf:resource="http://purl.org/dc/dcmitype/StillImage"/>
<dc:date>2026-09-09T18:32:39.096963</dc:date>
<dc:format>image/svg+xml</dc:format>
<dc:creator>
<cc:Agent>
<dc:title>Matplotlib v3.11.1, https://matplotlib.org/</dc:title>
</cc:Agent>
</dc:creator>
</cc:Work>
</rdf:RDF>
</metadata>
<defs>
<style type="text/css">*{stroke-linejoin: round; stroke-linecap: butt}</style>
</defs>
<g id="figure_1">
<g id="patch_1">
<path d="M 0 86.04
L 758.3677 86.04
L 758.3677 0
L 0 0
L 0 86.04
z
" style="fill: none"/>
</g>
<g id="axes_1">
<g id="patch_2">
<path d="M 1.44 80.82
L 38.571238 80.82
L 38.571238 67.958589
L 1.44 67.958589
z
" clip-path="url(#p4041b231bc)" style="fill: #6e6e6e"/>
</g>
<g id="patch_3">
<path d="M 1.44 65.135353
L 40.996285 65.135353
L 40.996285 52.273942
L 1.44 52.273942
z
" clip-path="url(#p4041b231bc)" style="fill: #6e6e6e"/>
</g>
<g id="patch_4">
<path d="M 1.44 49.450705
L 169.540418 49.450705
L 169.540418 36.589295
L 1.44 36.589295
z
" clip-path="url(#p4041b231bc)" style="fill: #6e6e6e"/>
</g>
<g id="patch_5">
<path d="M 1.44 33.766058
L 317.786457 33.766058
L 317.786457 20.904647
L 1.44 20.904647
z
" clip-path="url(#p4041b231bc)" style="fill: #6e6e6e"/>
</g>
<g id="patch_6">
<path d="M 1.44 18.081411
L 671.04 18.081411
L 671.04 5.22
L 1.44 5.22
z
" clip-path="url(#p4041b231bc)" style="fill: #2b6cb0"/>
</g>
<g id="text_1">
<!-- 3 100 user-agents -->
<g style="fill: #767676" transform="translate(46.606438 78.658335) scale(0.11 -0.11)">
<defs>
<path id="DejaVuSans-16" d="M 2597 2516
Q 3050 2419 3304 2112
Q 3559 1806 3559 1356
Q 3559 666 3084 287
Q 2609 -91 1734 -91
Q 1441 -91 1130 -33
Q 819 25 488 141
L 488 750
Q 750 597 1062 519
Q 1375 441 1716 441
Q 2309 441 2620 675
Q 2931 909 2931 1356
Q 2931 1769 2642 2001
Q 2353 2234 1838 2234
L 1294 2234
L 1294 2753
L 1863 2753
Q 2328 2753 2575 2939
Q 2822 3125 2822 3475
Q 2822 3834 2567 4026
Q 2313 4219 1838 4219
Q 1578 4219 1281 4162
Q 984 4106 628 3988
L 628 4550
Q 988 4650 1302 4700
Q 1616 4750 1894 4750
Q 2613 4750 3031 4423
Q 3450 4097 3450 3541
Q 3450 3153 3228 2886
Q 3006 2619 2597 2516
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-62" transform="scale(0.015625)"/>
<path id="DejaVuSans-14" d="M 794 531
L 1825 531
L 1825 4091
L 703 3866
L 703 4441
L 1819 4666
L 2450 4666
L 2450 531
L 3481 531
L 3481 0
L 794 0
L 794 531
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-13" d="M 2034 4250
Q 1547 4250 1301 3770
Q 1056 3291 1056 2328
Q 1056 1369 1301 889
Q 1547 409 2034 409
Q 2525 409 2770 889
Q 3016 1369 3016 2328
Q 3016 3291 2770 3770
Q 2525 4250 2034 4250
z
M 2034 4750
Q 2819 4750 3233 4129
Q 3647 3509 3647 2328
Q 3647 1150 3233 529
Q 2819 -91 2034 -91
Q 1250 -91 836 529
Q 422 1150 422 2328
Q 422 3509 836 4129
Q 1250 4750 2034 4750
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-3" transform="scale(0.015625)"/>
<path id="DejaVuSans-58" d="M 544 1381
L 544 3500
L 1119 3500
L 1119 1403
Q 1119 906 1312 657
Q 1506 409 1894 409
Q 2359 409 2629 706
Q 2900 1003 2900 1516
L 2900 3500
L 3475 3500
L 3475 0
L 2900 0
L 2900 538
Q 2691 219 2414 64
Q 2138 -91 1772 -91
Q 1169 -91 856 284
Q 544 659 544 1381
z
M 1991 3584
L 1991 3584
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-56" d="M 2834 3397
L 2834 2853
Q 2591 2978 2328 3040
Q 2066 3103 1784 3103
Q 1356 3103 1142 2972
Q 928 2841 928 2578
Q 928 2378 1081 2264
Q 1234 2150 1697 2047
L 1894 2003
Q 2506 1872 2764 1633
Q 3022 1394 3022 966
Q 3022 478 2636 193
Q 2250 -91 1575 -91
Q 1294 -91 989 -36
Q 684 19 347 128
L 347 722
Q 666 556 975 473
Q 1284 391 1588 391
Q 1994 391 2212 530
Q 2431 669 2431 922
Q 2431 1156 2273 1281
Q 2116 1406 1581 1522
L 1381 1569
Q 847 1681 609 1914
Q 372 2147 372 2553
Q 372 3047 722 3315
Q 1072 3584 1716 3584
Q 2034 3584 2315 3537
Q 2597 3491 2834 3397
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-48" d="M 3597 1894
L 3597 1613
L 953 1613
Q 991 1019 1311 708
Q 1631 397 2203 397
Q 2534 397 2845 478
Q 3156 559 3463 722
L 3463 178
Q 3153 47 2828 -22
Q 2503 -91 2169 -91
Q 1331 -91 842 396
Q 353 884 353 1716
Q 353 2575 817 3079
Q 1281 3584 2069 3584
Q 2775 3584 3186 3129
Q 3597 2675 3597 1894
z
M 3022 2063
Q 3016 2534 2758 2815
Q 2500 3097 2075 3097
Q 1594 3097 1305 2825
Q 1016 2553 972 2059
L 3022 2063
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-55" d="M 2631 2963
Q 2534 3019 2420 3045
Q 2306 3072 2169 3072
Q 1681 3072 1420 2755
Q 1159 2438 1159 1844
L 1159 0
L 581 0
L 581 3500
L 1159 3500
L 1159 2956
Q 1341 3275 1631 3429
Q 1922 3584 2338 3584
Q 2397 3584 2469 3576
Q 2541 3569 2628 3553
L 2631 2963
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-10" d="M 313 2009
L 1997 2009
L 1997 1497
L 313 1497
L 313 2009
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-44" d="M 2194 1759
Q 1497 1759 1228 1600
Q 959 1441 959 1056
Q 959 750 1161 570
Q 1363 391 1709 391
Q 2188 391 2477 730
Q 2766 1069 2766 1631
L 2766 1759
L 2194 1759
z
M 3341 1997
L 3341 0
L 2766 0
L 2766 531
Q 2569 213 2275 61
Q 1981 -91 1556 -91
Q 1019 -91 701 211
Q 384 513 384 1019
Q 384 1609 779 1909
Q 1175 2209 1959 2209
L 2766 2209
L 2766 2266
Q 2766 2663 2505 2880
Q 2244 3097 1772 3097
Q 1472 3097 1187 3025
Q 903 2953 641 2809
L 641 3341
Q 956 3463 1253 3523
Q 1550 3584 1831 3584
Q 2591 3584 2966 3190
Q 3341 2797 3341 1997
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-4a" d="M 2906 1791
Q 2906 2416 2648 2759
Q 2391 3103 1925 3103
Q 1463 3103 1205 2759
Q 947 2416 947 1791
Q 947 1169 1205 825
Q 1463 481 1925 481
Q 2391 481 2648 825
Q 2906 1169 2906 1791
z
M 3481 434
Q 3481 -459 3084 -895
Q 2688 -1331 1869 -1331
Q 1566 -1331 1297 -1286
Q 1028 -1241 775 -1147
L 775 -588
Q 1028 -725 1275 -790
Q 1522 -856 1778 -856
Q 2344 -856 2625 -561
Q 2906 -266 2906 331
L 2906 616
Q 2728 306 2450 153
Q 2172 0 1784 0
Q 1141 0 747 490
Q 353 981 353 1791
Q 353 2603 747 3093
Q 1141 3584 1784 3584
Q 2172 3584 2450 3431
Q 2728 3278 2906 2969
L 2906 3500
L 3481 3500
L 3481 434
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-51" d="M 3513 2113
L 3513 0
L 2938 0
L 2938 2094
Q 2938 2591 2744 2837
Q 2550 3084 2163 3084
Q 1697 3084 1428 2787
Q 1159 2491 1159 1978
L 1159 0
L 581 0
L 581 3500
L 1159 3500
L 1159 2956
Q 1366 3272 1645 3428
Q 1925 3584 2291 3584
Q 2894 3584 3203 3211
Q 3513 2838 3513 2113
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-57" d="M 1172 4494
L 1172 3500
L 2356 3500
L 2356 3053
L 1172 3053
L 1172 1153
Q 1172 725 1289 603
Q 1406 481 1766 481
L 2356 481
L 2356 0
L 1766 0
Q 1100 0 847 248
Q 594 497 594 1153
L 594 3053
L 172 3053
L 172 3500
L 594 3500
L 594 4494
L 1172 4494
z
" transform="scale(0.015625)"/>
</defs>
<use xlink:href="#DejaVuSans-16"/>
<use xlink:href="#DejaVuSans-62" transform="translate(63.625 0)"/>
<use xlink:href="#DejaVuSans-14" transform="translate(95.40625 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(159.03125 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(222.65625 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(286.28125 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(318.0625 0)"/>
<use xlink:href="#DejaVuSans-58" transform="translate(349.84375 0)"/>
<use xlink:href="#DejaVuSans-56" transform="translate(413.21875 0)"/>
<use xlink:href="#DejaVuSans-48" transform="translate(465.3125 0)"/>
<use xlink:href="#DejaVuSans-55" transform="translate(526.84375 0)"/>
<use xlink:href="#DejaVuSans-10" transform="translate(561.5625 0)"/>
<use xlink:href="#DejaVuSans-44" transform="translate(597.640625 0)"/>
<use xlink:href="#DejaVuSans-4a" transform="translate(658.921875 0)"/>
<use xlink:href="#DejaVuSans-48" transform="translate(722.40625 0)"/>
<use xlink:href="#DejaVuSans-51" transform="translate(783.9375 0)"/>
<use xlink:href="#DejaVuSans-57" transform="translate(847.3125 0)"/>
<use xlink:href="#DejaVuSans-56" transform="translate(886.515625 0)"/>
</g>
</g>
<g id="text_2">
<!-- 3 300 ua-parser -->
<g style="fill: #767676" transform="translate(49.031485 62.973687) scale(0.11 -0.11)">
<defs>
<path id="DejaVuSans-53" d="M 1159 525
L 1159 -1331
L 581 -1331
L 581 3500
L 1159 3500
L 1159 2969
Q 1341 3281 1617 3432
Q 1894 3584 2278 3584
Q 2916 3584 3314 3078
Q 3713 2572 3713 1747
Q 3713 922 3314 415
Q 2916 -91 2278 -91
Q 1894 -91 1617 61
Q 1341 213 1159 525
z
M 3116 1747
Q 3116 2381 2855 2742
Q 2594 3103 2138 3103
Q 1681 3103 1420 2742
Q 1159 2381 1159 1747
Q 1159 1113 1420 752
Q 1681 391 2138 391
Q 2594 391 2855 752
Q 3116 1113 3116 1747
z
" transform="scale(0.015625)"/>
</defs>
<use xlink:href="#DejaVuSans-16"/>
<use xlink:href="#DejaVuSans-62" transform="translate(63.625 0)"/>
<use xlink:href="#DejaVuSans-16" transform="translate(95.40625 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(159.03125 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(222.65625 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(286.28125 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(318.0625 0)"/>
<use xlink:href="#DejaVuSans-58" transform="translate(349.84375 0)"/>
<use xlink:href="#DejaVuSans-44" transform="translate(413.21875 0)"/>
<use xlink:href="#DejaVuSans-10" transform="translate(474.5 0)"/>
<use xlink:href="#DejaVuSans-53" transform="translate(510.578125 0)"/>
<use xlink:href="#DejaVuSans-44" transform="translate(574.0625 0)"/>
<use xlink:href="#DejaVuSans-55" transform="translate(635.34375 0)"/>
<use xlink:href="#DejaVuSans-56" transform="translate(676.453125 0)"/>
<use xlink:href="#DejaVuSans-48" transform="translate(728.546875 0)"/>
<use xlink:href="#DejaVuSans-55" transform="translate(790.078125 0)"/>
</g>
</g>
<g id="text_3">
<!-- 14 000 ua-parser (RE2) -->
<g style="fill: #767676" transform="translate(177.575618 47.28904) scale(0.11 -0.11)">
<defs>
<path id="DejaVuSans-17" d="M 2419 4116
L 825 1625
L 2419 1625
L 2419 4116
z
M 2253 4666
L 3047 4666
L 3047 1625
L 3713 1625
L 3713 1100
L 3047 1100
L 3047 0
L 2419 0
L 2419 1100
L 313 1100
L 313 1709
L 2253 4666
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-b" d="M 1984 4856
Q 1566 4138 1362 3434
Q 1159 2731 1159 2009
Q 1159 1288 1364 580
Q 1569 -128 1984 -844
L 1484 -844
Q 1016 -109 783 600
Q 550 1309 550 2009
Q 550 2706 781 3412
Q 1013 4119 1484 4856
L 1984 4856
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-35" d="M 2841 2188
Q 3044 2119 3236 1894
Q 3428 1669 3622 1275
L 4263 0
L 3584 0
L 2988 1197
Q 2756 1666 2539 1819
Q 2322 1972 1947 1972
L 1259 1972
L 1259 0
L 628 0
L 628 4666
L 2053 4666
Q 2853 4666 3247 4331
Q 3641 3997 3641 3322
Q 3641 2881 3436 2590
Q 3231 2300 2841 2188
z
M 1259 4147
L 1259 2491
L 2053 2491
Q 2509 2491 2742 2702
Q 2975 2913 2975 3322
Q 2975 3731 2742 3939
Q 2509 4147 2053 4147
L 1259 4147
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-28" d="M 628 4666
L 3578 4666
L 3578 4134
L 1259 4134
L 1259 2753
L 3481 2753
L 3481 2222
L 1259 2222
L 1259 531
L 3634 531
L 3634 0
L 628 0
L 628 4666
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-15" d="M 1228 531
L 3431 531
L 3431 0
L 469 0
L 469 531
Q 828 903 1448 1529
Q 2069 2156 2228 2338
Q 2531 2678 2651 2914
Q 2772 3150 2772 3378
Q 2772 3750 2511 3984
Q 2250 4219 1831 4219
Q 1534 4219 1204 4116
Q 875 4013 500 3803
L 500 4441
Q 881 4594 1212 4672
Q 1544 4750 1819 4750
Q 2544 4750 2975 4387
Q 3406 4025 3406 3419
Q 3406 3131 3298 2873
Q 3191 2616 2906 2266
Q 2828 2175 2409 1742
Q 1991 1309 1228 531
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-c" d="M 513 4856
L 1013 4856
Q 1481 4119 1714 3412
Q 1947 2706 1947 2009
Q 1947 1309 1714 600
Q 1481 -109 1013 -844
L 513 -844
Q 928 -128 1133 580
Q 1338 1288 1338 2009
Q 1338 2731 1133 3434
Q 928 4138 513 4856
z
" transform="scale(0.015625)"/>
</defs>
<use xlink:href="#DejaVuSans-14"/>
<use xlink:href="#DejaVuSans-17" transform="translate(63.625 0)"/>
<use xlink:href="#DejaVuSans-62" transform="translate(127.25 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(159.03125 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(222.65625 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(286.28125 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(349.90625 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(381.6875 0)"/>
<use xlink:href="#DejaVuSans-58" transform="translate(413.46875 0)"/>
<use xlink:href="#DejaVuSans-44" transform="translate(476.84375 0)"/>
<use xlink:href="#DejaVuSans-10" transform="translate(538.125 0)"/>
<use xlink:href="#DejaVuSans-53" transform="translate(574.203125 0)"/>
<use xlink:href="#DejaVuSans-44" transform="translate(637.6875 0)"/>
<use xlink:href="#DejaVuSans-55" transform="translate(698.96875 0)"/>
<use xlink:href="#DejaVuSans-56" transform="translate(740.078125 0)"/>
<use xlink:href="#DejaVuSans-48" transform="translate(792.171875 0)"/>
<use xlink:href="#DejaVuSans-55" transform="translate(853.703125 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(894.8125 0)"/>
<use xlink:href="#DejaVuSans-b" transform="translate(926.59375 0)"/>
<use xlink:href="#DejaVuSans-35" transform="translate(965.609375 0)"/>
<use xlink:href="#DejaVuSans-28" transform="translate(1035.09375 0)"/>
<use xlink:href="#DejaVuSans-15" transform="translate(1098.28125 0)"/>
<use xlink:href="#DejaVuSans-c" transform="translate(1161.90625 0)"/>
</g>
</g>
<g id="text_4">
<!-- 26 000 ua-parser (Rust) -->
<g style="fill: #767676" transform="translate(325.821657 31.604393) scale(0.11 -0.11)">
<defs>
<path id="DejaVuSans-19" d="M 2113 2584
Q 1688 2584 1439 2293
Q 1191 2003 1191 1497
Q 1191 994 1439 701
Q 1688 409 2113 409
Q 2538 409 2786 701
Q 3034 994 3034 1497
Q 3034 2003 2786 2293
Q 2538 2584 2113 2584
z
M 3366 4563
L 3366 3988
Q 3128 4100 2886 4159
Q 2644 4219 2406 4219
Q 1781 4219 1451 3797
Q 1122 3375 1075 2522
Q 1259 2794 1537 2939
Q 1816 3084 2150 3084
Q 2853 3084 3261 2657
Q 3669 2231 3669 1497
Q 3669 778 3244 343
Q 2819 -91 2113 -91
Q 1303 -91 875 529
Q 447 1150 447 2328
Q 447 3434 972 4092
Q 1497 4750 2381 4750
Q 2619 4750 2861 4703
Q 3103 4656 3366 4563
z
" transform="scale(0.015625)"/>
</defs>
<use xlink:href="#DejaVuSans-15"/>
<use xlink:href="#DejaVuSans-19" transform="translate(63.625 0)"/>
<use xlink:href="#DejaVuSans-62" transform="translate(127.25 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(159.03125 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(222.65625 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(286.28125 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(349.90625 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(381.6875 0)"/>
<use xlink:href="#DejaVuSans-58" transform="translate(413.46875 0)"/>
<use xlink:href="#DejaVuSans-44" transform="translate(476.84375 0)"/>
<use xlink:href="#DejaVuSans-10" transform="translate(538.125 0)"/>
<use xlink:href="#DejaVuSans-53" transform="translate(574.203125 0)"/>
<use xlink:href="#DejaVuSans-44" transform="translate(637.6875 0)"/>
<use xlink:href="#DejaVuSans-55" transform="translate(698.96875 0)"/>
<use xlink:href="#DejaVuSans-56" transform="translate(740.078125 0)"/>
<use xlink:href="#DejaVuSans-48" transform="translate(792.171875 0)"/>
<use xlink:href="#DejaVuSans-55" transform="translate(853.703125 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(894.8125 0)"/>
<use xlink:href="#DejaVuSans-b" transform="translate(926.59375 0)"/>
<use xlink:href="#DejaVuSans-35" transform="translate(965.609375 0)"/>
<use xlink:href="#DejaVuSans-58" transform="translate(1030.609375 0)"/>
<use xlink:href="#DejaVuSans-56" transform="translate(1093.984375 0)"/>
<use xlink:href="#DejaVuSans-57" transform="translate(1146.078125 0)"/>
<use xlink:href="#DejaVuSans-c" transform="translate(1185.28125 0)"/>
</g>
</g>
<g id="text_5">
<!-- 56 000 uarite -->
<g style="fill: #767676" transform="translate(679.0752 15.920175) scale(0.11 -0.11)">
<defs>
<path id="DejaVuSans-18" d="M 691 4666
L 3169 4666
L 3169 4134
L 1269 4134
L 1269 2991
Q 1406 3038 1543 3061
Q 1681 3084 1819 3084
Q 2600 3084 3056 2656
Q 3513 2228 3513 1497
Q 3513 744 3044 326
Q 2575 -91 1722 -91
Q 1428 -91 1123 -41
Q 819 9 494 109
L 494 744
Q 775 591 1075 516
Q 1375 441 1709 441
Q 2250 441 2565 725
Q 2881 1009 2881 1497
Q 2881 1984 2565 2268
Q 2250 2553 1709 2553
Q 1456 2553 1204 2497
Q 953 2441 691 2322
L 691 4666
z
" transform="scale(0.015625)"/>
<path id="DejaVuSans-4c" d="M 603 3500
L 1178 3500
L 1178 0
L 603 0
L 603 3500
z
M 603 4863
L 1178 4863
L 1178 4134
L 603 4134
L 603 4863
z
" transform="scale(0.015625)"/>
</defs>
<use xlink:href="#DejaVuSans-18"/>
<use xlink:href="#DejaVuSans-19" transform="translate(63.625 0)"/>
<use xlink:href="#DejaVuSans-62" transform="translate(127.25 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(159.03125 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(222.65625 0)"/>
<use xlink:href="#DejaVuSans-13" transform="translate(286.28125 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(349.90625 0)"/>
<use xlink:href="#DejaVuSans-3" transform="translate(381.6875 0)"/>
<use xlink:href="#DejaVuSans-58" transform="translate(413.46875 0)"/>
<use xlink:href="#DejaVuSans-44" transform="translate(476.84375 0)"/>
<use xlink:href="#DejaVuSans-55" transform="translate(538.125 0)"/>
<use xlink:href="#DejaVuSans-4c" transform="translate(579.234375 0)"/>
<use xlink:href="#DejaVuSans-57" transform="translate(607.015625 0)"/>
<use xlink:href="#DejaVuSans-48" transform="translate(646.21875 0)"/>
</g>
</g>
</g>
</g>
<defs>
<clipPath id="p4041b231bc">
<rect x="1.44" y="1.44" width="669.6" height="83.16"/>
</clipPath>
</defs>
</svg>

After

Width:  |  Height:  |  Size: 19 KiB

+2 -3
View File
@@ -8,19 +8,18 @@ build-backend = "hatchling.build"
[project]
name = "uarite"
dynamic = ["version"]
description = "Compact, human-readable User-Agent formatting and bot detection — no regex database, stdlib only"
description = "User-Agent parsing done right. Accurate, small, pure Python and fast."
authors = [
{ name = "Leo Vasanko" },
]
keywords = ["user-agent", "ua-parser"]
readme = "README.md"
license = "MIT OR Unlicense"
requires-python = ">=3.10"
classifiers = [
"Intended Audience :: Developers",
"Operating System :: OS Independent",
"Programming Language :: Python :: 3",
"Topic :: Internet :: WWW/HTTP",
"Topic :: Software Development :: Libraries :: Python Modules",
]
dependencies = []
+2 -2
View File
@@ -17,7 +17,7 @@ benchmark-only dependencies, pulled ad hoc via `uv run --with`.
- `prettytable.py` — prints the accuracy-comparison markdown table:
```
uv run --with ua-parser --with user-agents python scripts/prettytable.py
uv run --with ua-parser --with fastuaparser python scripts/prettytable.py
```
- `bench.py` — prints browser-accuracy counts, crawler-detection rates,
@@ -25,6 +25,6 @@ benchmark-only dependencies, pulled ad hoc via `uv run --with`.
mix, bot storm) with cache statistics:
```
uv run --with ua-parser --with user-agents --with user-agent-parser \
uv run --with ua-parser --with user-agents --with fastuaparser \
python scripts/bench.py
```
+141 -51
View File
@@ -1,14 +1,25 @@
"""Benchmark uarite vs ua-parser vs user-agents vs user-agent-parser.
# /// script
# requires-python = ">=3.14"
# dependencies = [
# "fastuaparser>=0.1.4",
# "ua-parser[re2,regex]>=1.0.2",
# "uarite",
# "user-agents>=2.2.0",
# ]
#
# [tool.uv.sources]
# uarite = { path = "..", editable = true }
# ///
"""Benchmark uarite vs ua-parser vs user-agents vs fastuaparser.
Reproduces the README's numbers: browser accuracy on 100 modern UAs,
crawler detection on 2163 real-world crawler UAs, and timing (unique UAs,
a realistic repeat/unique mix, a pure bot storm) with cache introspection.
Data lives in scripts/data (see download_data.py). Requires uarite
(installed) plus the benchmark-only reference parsers:
Data lives in scripts/data (see download_data.py). All dependencies,
including uarite itself (editable), are declared inline:
uv run --with ua-parser --with user-agents --with user-agent-parser \
python scripts/bench.py
uv run scripts/bench.py
"""
import json
@@ -17,17 +28,56 @@ import re
import timeit
from pathlib import Path
import ua_parser
from fastuaparser import parse_ua as fua_parse
from ua_parser import parse as ua_parse
from user_agent_parser import parse as uap_parse
from user_agents import parse as uas_parse
from uarite import uaparse
from uarite.core import _parse_client
ALL_DOMAINS = (
ua_parser.Domain.USER_AGENT | ua_parser.Domain.OS | ua_parser.Domain.DEVICE
)
_VARIANTS = {}
def ua_variant(name):
"""ua-parser with a specific resolver backend (pure/re2/rust), lazily
built so its one-time database load lands in the untimed warm-up call.
The default parse() picks whichever native backend is installed, so
backends must be forced explicitly to benchmark them separately."""
if name not in _VARIANTS:
ctor = {
"pure": ua_parser.BasicResolver,
"re2": ua_parser.Re2Resolver,
"rust": ua_parser.RegexResolver,
}[name]
parser = ua_parser.Parser(
ua_parser.CachingResolver(
ctor(ua_parser.load_builtins()), ua_parser.Cache(2000)
)
)
_VARIANTS[name] = lambda ua: parser(ua, ALL_DOMAINS)
return _VARIANTS[name]
def uap_pure(ua):
return ua_variant("pure")(ua)
def uap_re2(ua):
return ua_variant("re2")(ua)
def uap_rust(ua):
return ua_variant("rust")(ua)
DATA = Path(__file__).parent / "data"
BROWSERS = json.loads((DATA / "top-user-agents.json").read_text())
CRAWLERS = json.loads((DATA / "crawler-user-agents.json").read_text())
OWN = (DATA / "ua.txt").read_text().splitlines()
CRAWLER_UAS = [ua for c in CRAWLERS for ua in (c.get("instances") or [c["pattern"]])]
@@ -76,7 +126,6 @@ def score_browsers():
for name, fn in (
("ua-parser", ua_parse),
("user-agents", uas_parse),
("user-agent-parser", uap_parse),
):
fam_ok = ver_ok = os_ok = 0
for ua in BROWSERS:
@@ -86,10 +135,6 @@ def score_browsers():
fam = r.user_agent.family or ""
maj = r.user_agent.major or ""
osf = r.os.family or ""
elif name == "user-agent-parser":
fam = r[0] or ""
maj = (r[1] or "").split(".")[0]
osf = r[2] or ""
else:
fam = r.browser.family or ""
maj = str(r.browser.version[0]) if r.browser.version else ""
@@ -99,15 +144,44 @@ def score_browsers():
ver_ok += emaj == maj
os_ok += eos == norm_os(osf)
res[name] = (fam_ok, ver_ok, os_ok)
# fastuaparser returns one pretty string ("Chrome - Windows") and no
# version numbers, so score it on the substrings it does produce.
fam_ok = ver_ok = os_ok = 0
for ua in BROWSERS:
efam, emaj, eos = expected(ua)
pretty = fua_parse(ua)
fam_ok += efam.split()[0].lower() in pretty.lower()
ver_ok += bool(emaj) and emaj in pretty
os_ok += bool(
re.search(
{
"ios": r"iOS|iPhone|iPad",
"macos": r"Mac",
"windows": r"Windows",
"linux": r"Linux",
"android": r"Android",
}[eos],
pretty,
)
if eos
else True
)
res["fastuaparser"] = (fam_ok, ver_ok, os_ok)
fam_ok = ver_ok = os_ok = 0
for ua in BROWSERS:
efam, emaj, eos = expected(ua)
r = uaparse(ua)
if emaj:
if eos == "ios":
# Safari is the only browser iOS has, so the right answer is the
# device and the iOS version: "iPhone iOS 17", not "Safari/17".
device = "iPhone" if "iPhone" in ua else "iPad"
m = re.search(r"OS (\d+)", ua)
fam_ok += r.pretty.startswith(device)
ver_ok += bool(m and m.group(1) in r.pretty)
elif emaj:
fam_ok += f"{efam}/{emaj}" in r.pretty
ver_ok += f"/{emaj}" in r.pretty
else:
# iOS webview: no version; "iPhone iOS 18" is the right answer
fam_ok += efam in r.pretty or "iPhone" in r.pretty or "iPad" in r.pretty
ver_ok += 1
oslabel = {
@@ -128,7 +202,7 @@ def score_browsers():
def score_crawlers():
out = {}
crashes = 0
for name in ("ua-parser", "user-agents", "user-agent-parser", "uarite"):
for name in ("ua-parser", "user-agents", "fastuaparser", "uarite"):
det = 0
for ua in CRAWLER_UAS:
try:
@@ -140,12 +214,16 @@ def score_crawlers():
)
elif name == "user-agents":
bot = uas_parse(ua).is_bot
elif name == "user-agent-parser":
bot = uap_parse(ua)[4] == "Bot"
elif name == "fastuaparser":
# "Bot" is a real verdict; "Other" is a non-browser
# client (wget etc.), which is automated traffic too.
bot = fua_parse(ua).split(" - ")[0] in ("Bot", "Other")
else:
bot = bool(uaparse(ua).bot)
# Anything not recognized as a real browser is automated:
# known bots, generic spiders, clients, spoofed claims.
bot = uaparse(ua).kind != "browser"
except Exception:
crashes += name == "user-agent-parser"
crashes += name == "fastuaparser"
continue
det += bot
out[name] = det
@@ -187,9 +265,11 @@ def bench_realistic():
f" with 2000 mostly-unique bots)"
)
for name, fn in (
("ua-parser", ua_parse),
("ua-parser", uap_pure),
("ua-parser (re2)", uap_re2),
("ua-parser (rust)", uap_rust),
("user-agents", uas_parse),
("user-agent-parser", uap_parse),
("fastuaparser", fua_parse),
("uarite", uaparse),
):
fn = safe(fn)
@@ -198,9 +278,11 @@ def bench_realistic():
print(f"{name:20} {t / len(mix) * 1e6:7.1f} µs/UA cache: {info}")
print(f"\n## pure bot storm ({len(storm)} unique UAs, zero cache value)")
for name, fn in (
("ua-parser", ua_parse),
("ua-parser", uap_pure),
("ua-parser (re2)", uap_re2),
("ua-parser (rust)", uap_rust),
("user-agents", uas_parse),
("user-agent-parser", uap_parse),
("fastuaparser", fua_parse),
("uarite", uaparse),
):
fn = safe(fn)
@@ -222,56 +304,64 @@ def safe(fn):
def cache_info(name):
if name == "uarite":
i = _parse_client.cache_info()
return f"{i.hits} hits / {i.misses} misses (cap 1024, browsers only)"
if name == "user-agent-parser":
from user_agent_parser.parser import _cached_parse_user_agent
i = _cached_parse_user_agent.cache_info()
return f"{i.hits} hits / {i.misses} misses (cap 512 LRU)"
i = uaparse.cache_info()
return f"{i.hits} hits / {i.misses} misses (cap 1024)"
if name == "fastuaparser":
return "none (branch parser, ~1 µs flat)"
if name == "user-agents":
from ua_parser.user_agent_parser import _PARSE_CACHE
return f"{len(_PARSE_CACHE)} entries (cap 200, CLEARS when full)"
if name == "ua-parser":
if name.startswith("ua-parser"):
return "cap 2000 S3-FIFO (scan-resistant)"
return ""
def bench():
alluas = BROWSERS + CRAWLER_UAS + OWN
n = 3
"""Cold-cache speed: a single pass over previously unseen UAs, with
equal shares of realistic browser and crawler strings since they take
different parse paths. Runs before the accuracy passes, which would
otherwise warm every parser's cache with these very strings.
Each parser first parses one dummy UA (untimed) so that lazy regex
compilation and database loading do not land on the first real item —
ua-parser's first parse alone costs ~59 ms loading its database. The
cache gains nothing from it since all timed UAs are unique."""
rng = random.Random(7)
work = BROWSERS + rng.sample(CRAWLER_UAS, len(BROWSERS))
rng.shuffle(work)
res = {}
for name, fn in (
("ua-parser", ua_parse),
("ua-parser", uap_pure),
("ua-parser (re2)", uap_re2),
("ua-parser (rust)", uap_rust),
("user-agents", uas_parse),
("user-agent-parser", uap_parse),
("fastuaparser", fua_parse),
("uarite", uaparse),
):
fn = safe(fn)
t = timeit.timeit(lambda: [fn(u) for u in alluas], number=n)
res[name] = t / n / len(alluas) * 1e6
# warm cache: repeat a small realistic working set many times
working = (BROWSERS + OWN[:50]) * 10
t = timeit.timeit(lambda: [uaparse(u) for u in working], number=n)
res["uarite (warm cache)"] = t / n / len(working) * 1e6
return res
fn("Warmup/1.0 (+https://example.com/warmup)")
uaparse.cache_clear()
t = timeit.timeit(lambda: [fn(u) for u in work], number=1)
res[name] = t / len(work) * 1e6
return res, len(work)
if __name__ == "__main__":
print(f"## browser accuracy (n={len(BROWSERS)}): family / version / OS correct")
res, nwork = bench()
print(
f"## speed (µs per cold parse, {nwork} unique UAs,"
" half browsers / half crawlers)"
)
for k, v in res.items():
print(f"{k:20} {v:8.1f}")
print(f"\n## browser accuracy (n={len(BROWSERS)}): family / version / OS correct")
for k, (f, v, o) in score_browsers().items():
print(f"{k:14} {f:3}/100 {v:3}/100 {o:3}/100")
det, url_have, url_got, crashes = score_crawlers()
print(f"\n## crawler detection (n={len(CRAWLER_UAS)})")
for k, v in det.items():
print(f"{k:18} {v:5} ({v / len(CRAWLER_UAS):.1%})")
print(f"user-agent-parser crashed on {crashes} UAs")
print(f"fastuaparser crashed on {crashes} UAs")
print(f"\nuarite URL extraction: {url_got}/{url_have} of instances carrying a URL")
print(
"\n## speed (µs per parse, mixed set of %d UAs)"
% (len(BROWSERS) + len(CRAWLER_UAS) + len(OWN))
)
for k, v in bench().items():
print(f"{k:20} {v:8.1f}")
bench_realistic()
+68
View File
@@ -0,0 +1,68 @@
"""Generate the JS data tables for uarite-js from the Python source.
Run from the repository root:
uv run scripts/port_tables.py
Writes uarite-js/src/tables.ts. That file is generated — edit the Python
tables (uarite/bots.py, uarite/clients.py) instead and re-run this script.
"""
import json
from pathlib import Path
from uarite.bots import BOTS, KIND_LABEL, LABELED, PRETTY_OVERRIDE, PROVIDER_OF
from uarite.clients import BROWSERS, SAMSUNG, SAMSUNG_SERIES
OUT = Path(__file__).resolve().parent.parent / "uarite-js" / "src" / "tables.ts"
HEADER = """\
// Generated by scripts/port_tables.py from uarite/bots.py and
// uarite/clients.py. Do not edit by hand; re-run the script.
"""
def js(obj: object) -> str:
return json.dumps(obj, indent=2, ensure_ascii=False)
def main() -> None:
src = [HEADER]
src.append(
"/** UA token -> display name, checked in order; first hit wins. */\n"
f"export const BROWSERS: ReadonlyArray<readonly [string, string]> = {js(list(map(list, BROWSERS)))};\n"
)
src.append(
"/** Lowercase UA substring -> [display name, kind]. First match wins. */\n"
f"export const BOTS: Readonly<Record<string, readonly [string, string]>> = {js({k: list(v) for k, v in BOTS.items()})};\n"
)
src.append(
"/** Pretty suffixes for the kinds more precise than a generic spider. */\n"
f"export const KIND_LABEL: Readonly<Record<string, string>> = {js(KIND_LABEL)};\n"
)
src.append(
"/** Bot display name -> provider. */\n"
f"export const PROVIDER_OF: Readonly<Record<string, string>> = {js(PROVIDER_OF)};\n"
)
src.append(
"/** Per-bot pretty overrides: the full display string. */\n"
f"export const PRETTY_OVERRIDE: Readonly<Record<string, string>> = {js(PRETTY_OVERRIDE)};\n"
)
src.append(
"/** Bot names whose kind label is displayed. */\n"
f"export const LABELED: ReadonlySet<string> = new Set({js(sorted(LABELED))});\n"
)
src.append(
"/** Samsung model code (without region/carrier letter) -> marketing name. */\n"
f"export const SAMSUNG: Readonly<Record<string, string>> = {js(SAMSUNG)};\n"
)
src.append(
"/** Samsung series for codes missing from SAMSUNG. */\n"
f"export const SAMSUNG_SERIES: Readonly<Record<string, string>> = {js(SAMSUNG_SERIES)};\n"
)
OUT.write_text("\n".join(src))
print(f"wrote {OUT}")
if __name__ == "__main__":
main()
+8 -19
View File
@@ -1,15 +1,15 @@
"""Print the README's accuracy-comparison table as markdown.
uarite's column is simply `r.pretty`. The reference modules have no
display format; their columns use one plain format string each on their
structured output (footnotes ²³ in the README). Requires uarite (installed)
plus the benchmark-only reference parsers:
uarite's and fastuaparser's columns are simply their pretty-string output.
ua-parser has no display format; its column uses one plain format string
on the structured output (footnotes ¹² in the README). Requires uarite
(installed) plus the benchmark-only reference parsers:
uv run --with ua-parser --with user-agents python scripts/prettytable.py
uv run --with ua-parser --with fastuaparser python scripts/prettytable.py
"""
from fastuaparser import parse_ua as fua_parse
from ua_parser import parse as ua_parse
from user_agents import parse as uas_parse
from uarite import uaparse
@@ -55,10 +55,6 @@ CASES = [
"Googlebot (disguised)",
"Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.84 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
),
(
"Claude-SearchBot (disguised)",
"Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/148.0.3245.171 Mobile Safari/537.36; compatible; Claude-SearchBot/1.0; +https://www.anthropic.com/claude-searchbot",
),
(
"Facebook preview (disguised)",
"Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/134.0.6885.65 Mobile Safari/537.36; compatible; facebookexternalhit/1.1; +http://www.facebook.com/externalhit_uatext.php",
@@ -90,15 +86,8 @@ def imitate_uaparser(ua):
return f"{fam}/{maj} {osf} {dev}"
def imitate_useragents(ua):
r = uas_parse(ua)
fam = r.browser.family or ""
maj = str(r.browser.version[0]) if r.browser.version else ""
return f"{fam}/{maj} {r.os.family or ''} {r.device.family or ''}"
print("| Case | uarite¹ | ua-parser² | user-agents³ |")
print("| Case | uarite¹ | fastuaparser² | ua-parser³ |")
print("|---|---|---|---|")
for name, ua in CASES:
ours = uaparse(ua).pretty or ""
print(f"| {name} | {ours} | {imitate_uaparser(ua)} | {imitate_useragents(ua)} |")
print(f"| {name} | {ours} | {fua_parse(ua)} | {imitate_uaparser(ua)} |")
+67
View File
@@ -0,0 +1,67 @@
# /// script
# requires-python = ">=3.14"
# dependencies = ["matplotlib"]
# ///
"""Bar chart of cold-parse throughput for the README's Performance section.
Numbers are pasted from `uv run scripts/bench.py` (the cold 50/50
browser/crawler mix, µs per parse) and shown as parses per second.
Transparent SVG, text baked to paths, neutral grays: renders the same
on light and dark themes. All labels sit on the bars themselves.
Very wide aspect ratio: forges render images at full content width,
so height alone controls how tall it appears.
uv run scripts/speedplot.py
"""
from pathlib import Path
import matplotlib
matplotlib.use("Agg")
matplotlib.rcParams["svg.fonttype"] = "path" # text as paths: renders anywhere
import matplotlib.pyplot as plt # noqa: E402
# µs per cold parse, from scripts/bench.py.
# fastuaparser (~1 µs, 1M parses/s) and cached results are too far off
# this scale to draw meaningfully and are only mentioned in the README.
US = {
"uarite": 18.0,
"ua-parser (Rust)": 38.1,
"ua-parser (RE2)": 71.7,
"ua-parser": 304.7,
"user-agents": 324.6,
}
#: Readable on both white and dark backgrounds.
OUTSIDE = "#767676"
data = sorted(((n, 1e6 / us) for n, us in US.items()), key=lambda t: -t[1])
names = [n for n, _ in data][::-1]
values = [v for _, v in data][::-1]
colors = ["#6e6e6e"] * len(data)
colors[names.index("uarite")] = "#2b6cb0"
fig, ax = plt.subplots(figsize=(12, 1.5), dpi=100)
bars = ax.barh(names, values, color=colors, height=0.82)
xmax = max(values)
ax.set_xlim(0, xmax)
ax.axis("off")
for bar, name, v in zip(bars, names, values):
# Round to two significant digits: 121951 -> "120 000".
rounded = round(v, 1 - int(f"{v:.0e}".split("e")[1]))
label = f"{rounded:,.0f} {name}".replace(",", " ")
# va="center" centers the font bbox incl. descender space, which leaves
# the glyphs slightly high; nudge down to optically center on the bar.
y = bar.get_y() + bar.get_height() / 2 - 0.09
# All labels after the bar: theme-neutral gray, number before name.
# The longest bar's label overflows the axes; bbox_inches="tight"
# below expands the canvas to include it, so no dead space remains.
ax.text(bar.get_width() + xmax * 0.012, y, label,
va="center", color=OUTSIDE, fontsize=11)
out = Path("docs/bench-speed.svg")
out.parent.mkdir(exist_ok=True)
fig.savefig(out, transparent=True, bbox_inches="tight", pad_inches=0.02)
print("wrote", out)
+3
View File
@@ -0,0 +1,3 @@
{
"semi": false
}
+89
View File
@@ -0,0 +1,89 @@
# User-Agent Parsing Done Right
Fast and accurate handling of modern browsers and crawlers. Despite its light weight and no dependencies, uarite identifies both browsers and crawlers more accurately than any competing implementation tested here. We also provide a [Python uarite](https://pypi.org/project/uarite/) with exact same output.
It returns structured classification, but also the thing most applications eventually need: **a short pretty description**.
## Usage
Add it to your project:
```sh
npm install @vasanko/uarite
```
```js
import { uaparse } from "@vasanko/uarite"
const { pretty, engine, os, kind } = uaparse(
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36",
)
// Chrome/152 Windows, Chromium, Windows, browser
const { pretty, kind, url, provider } = uaparse(
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot",
)
// GPTBot (AI), ai, https://openai.com/gptbot, OpenAI
```
Plain HTML? A prebuilt minified ESM bundle you can host yourself or link from CDN:
```html
<script type="module">
import { uaparse } from "https://cdn.jsdelivr.net/npm/@vasanko/uarite/dist/uarite.min.js"
console.log(uaparse(navigator.userAgent))
</script>
```
## Output
`uaparse(ua)` returns a `UA` object with string fields. Any field may be empty string when the information is unavailable.
| Field | Content |
| -------- | ------------------------------------------------ |
| pretty | Compact display string; raw UA when unrecognized |
| engine | Chromium, Gecko, Safari, ArkWeb |
| os | Windows, macOS, Linux, iOS, Android, HarmonyOS |
| kind | browser, ai, search, social, analytics, spider |
| url | Crawler information URL |
| provider | Provider of a known crawler family |
The pretty field is intended for UIs and logs. The url can be attached to it as a link when available.
The engine and os fields are intentionally broad. The kind field distinguishes browsers from AI collectors, search engines, social previews, monitoring tools, generic spiders, and ordinary HTTP clients. Any non-browser kind represents automated traffic.
Detection is necessarily limited by what the User-Agent reveals. Crawlers can masquerade as ordinary browsers or other crawlers, so sites that need stronger identification should use additional methods rather than relying on UA detection alone.
## Comparison
The popular npm options for this task are bowser and ua-parser-js. The table below compares representative User-Agent formats.
| Case | uarite¹ | ua-parser-js² | bowser³ |
| ---------------------------- | ------------------------- | ------------------------------------ | ------------------------------------ |
| Chrome, Windows | Chrome/152 Windows | Chrome/152 Windows | Chrome/152.0.0.0 Windows |
| Chrome, Android (no model) | Chrome/152 Android | Mobile Chrome/152 Android K❌ | Chrome/152.0.0.0 Android |
| Edge, Android (model code) | Edge/110 Galaxy S7 | Edge/110 Android SM-G930P | Microsoft Edge/110.0.1587.66 Android |
| Safari, iPhone | iPhone iOS 17 | Mobile Safari/17 iOS iPhone | Safari/17.0 iOS iPhone |
| Huawei HarmonyOS phone | HuaweiBrowser/6 HarmonyOS | Huawei Browser/6 HarmonyOS ALN-AL00 | Android Browser/ Android ❌ |
| GPTBot | GPTBot (AI) | WebKit/537 ❌ | GPTBot/1.2 |
| Googlebot (disguised) | Googlebot (search) | Mobile Chrome/122 Android Nexus 5 ❌ | Googlebot/2.1 Android❌ |
| Facebook preview (disguised) | Facebook | Mobile Chrome/134 Android Pixel 7 ❌ | FacebookExternalHit/ Android❌ |
| WhatsApp preview | WhatsApp | (nothing) ❌ | WhatsApp/2.23.20.0 |
| python-requests | python-requests/2.32.5 | (nothing) ❌ | (nothing) ❌ |
- ❌ marks incorrect data such as an OS from a crawler's disguise, a frozen compat placeholder reported as a device, or a missed identity
- ¹ `uaparse(ua).pretty` shown as is
- ² `{browser.name??''}/{browser.major??''} {os.name??''} {device.model??''}`; free MIT tier of v2
- ³ `{browser.name??''}/{browser.version??''} {os.name??''} {platform.model??''}`
Parsing a mixed set of 316 real-world browser and crawler UAs, parses per second: **uarite 580 000**, bowser 130 000 and ua-parser-js 18 000. The other parsers don't appear to implement caching. For previously seen UA strings, however, uarite reaches **5.6 million**, while all others remain at the rates quoted above.
All are relatively small: **uarite minifies to 9 kB**, bowser to 37 kB and ua-parser-js to 28 kB.
## Why yet another UA parser
Rather than relying on a large historical regex database, the parser focuses on modern UA formats and parses them directly, choosing the most specific interpretation available. This keeps the implementation small while handling today's browsers and crawler traffic well. Until now, I had been using those other modules and building my own pretty-UA formatting on top of them, fixing by post processing issues the upstream didn't care of.
Eventually it became easier to start over with a parser designed around modern traffic. The result is uarite.
Hopefully it helps you too. Star my [GitHub](https://github.com/leovasanko/uarite) if it did.
+1
View File
File diff suppressed because one or more lines are too long
+43
View File
@@ -0,0 +1,43 @@
{
"name": "@vasanko/uarite",
"version": "0.2.2",
"description": "User-Agent parsing done right. Accurate, small, dependency-free and fast.",
"keywords": [
"user-agent",
"ua-parser"
],
"homepage": "https://git.zi.fi/LeoVasanko/uarite",
"repository": "https://git.zi.fi/LeoVasanko/uarite",
"bugs": "https://github.com/LeoVasanko/uarite",
"license": "MIT OR Unlicense",
"publishConfig": {
"access": "public"
},
"type": "module",
"exports": {
".": {
"types": "./dist/index.d.ts",
"default": "./dist/index.js"
}
},
"files": [
"dist"
],
"engines": {
"node": ">=18"
},
"scripts": {
"build": "tsc && npm run build:min",
"build:min": "esbuild src/index.ts --bundle --minify --format=esm --target=es2022 --outfile=dist/uarite.min.js",
"format": "prettier --write src test",
"format:check": "prettier --check src test",
"gen:tables": "cd .. && uv run scripts/port_tables.py && cd uarite-js && prettier --write src/tables.ts",
"test": "npm run build && node --test test/*.test.js",
"prepublishOnly": "npm run build && npm test"
},
"devDependencies": {
"esbuild": "^0.28.2",
"prettier": "^3.9.6",
"typescript": "^5.6.0"
}
}
+311
View File
@@ -0,0 +1,311 @@
import {
BOTS,
BROWSERS,
KIND_LABEL,
LABELED,
PRETTY_OVERRIDE,
PROVIDER_OF,
SAMSUNG,
SAMSUNG_SERIES,
} from "./tables.js"
/** Result of parsing a User-Agent string. Any field may be empty. */
export interface UA {
/** Compact display string; raw UA when unrecognized. */
pretty: string
/** Chromium, Gecko, Safari, ArkWeb. */
engine: string
/** Windows, macOS, Linux, iOS, Android, HarmonyOS. */
os: string
/** browser, ai, search, social, analytics, spider. */
kind: string
/** Crawler information URL. */
url: string
/** Provider of a known crawler family. */
provider: string
}
const EMPTY: UA = {
pretty: "",
engine: "",
os: "",
kind: "",
url: "",
provider: "",
}
/** Fallback for unknown crawlers: a product token whose name says so. */
const BOT_TOKEN = /[^\s();/]*(?:bot|spider|crawl|scan|verif|check)[^\s();/]*/i
/** A "+https://…" pointer is a crawler tell; real browsers carry no URL. */
const URL = /\+\s*(https?:\/\/[^\s;)]+)/
const ANY_URL = /(https?:\/\/[^\s;)]+)/
/** Crawler name next to the info URL: "compatible; Page2RSS/0.7; +http://…". */
const COMPATIBLE_NAME = /compatible;\s*([^;/()]+?)(?:\/[\d.vx]+)?\s*;/
/** Name from an info URL's host when no product token is available. */
const HOST = /https?:\/\/(?:www\.)?([^/\s;)]+)/
/** Android model token: "Android 15; SM-S918B)", "Android 12; Pixel 6; Build/…". */
const ANDROID_MODEL = /Android [\d.]+; ([^;()]+?)(?:;|\)| Build\/)/
/** All known-bot substrings in one compiled pass. */
const BOTS_RE = new RegExp(
Object.keys(BOTS)
.map((s) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"))
.join("|"),
)
/** The crawler's info URL from the UA, or "" (mailto: is not a URL). */
export function url(ua: string): string {
const m = URL.exec(ua) ?? ANY_URL.exec(ua)
return m ? m[1] : ""
}
/** (display name, kind) of the crawler/unfurler the UA claims. */
export function bot(ua: string): readonly [string, string] {
const low = ua.toLowerCase()
const m = BOTS_RE.exec(low)
if (m) return BOTS[m[0]]
// Cheap keyword gates keep the regexes off the hot path.
if (/(?:bot|spider|crawl|scan|verif|check)/.test(low)) {
const t = BOT_TOKEN.exec(ua)
if (t) {
const name = t[0].replace(/^;+|;+$/g, "")
return [name[0].toUpperCase() + name.slice(1), "spider"]
}
}
// An info URL in the UA is a crawler convention. Name it from the
// "compatible; Name/x" token, else the first product token, else the
// URL's host.
if (ua.includes("://")) {
let name = ""
const cm = COMPATIBLE_NAME.exec(ua)
if (cm) {
name = cm[1].trim()
} else if (!ua.startsWith("Mozilla")) {
name = ua.split(" ")[0].split("/")[0]
}
const nl = name.toLowerCase()
if (!name || nl.startsWith("mozilla") || nl.startsWith("msie")) {
const hm = HOST.exec(ua)
name = hm ? hm[1] : ""
}
if (name) return [name[0].toUpperCase() + name.slice(1), "spider"]
}
return ["", ""]
}
/** Major version of a `token/x.y` product in the UA, or "". */
export function version(ua: string, token: string): string {
const m = new RegExp(
`${token.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")}/(\\d+)`,
).exec(ua)
return m ? m[1] : ""
}
/** `Browser/major` for the browsers we care to distinguish. */
export function browser(ua: string): string {
for (const [token, name] of BROWSERS) {
const ver = version(ua, token)
if (ver) return `${name}/${ver}`
}
if (ua.includes("Safari/")) {
const ver = version(ua, "Version")
return ver ? `Safari/${ver}` : "Safari"
}
return ""
}
/** Engines that differ from the Chromium default for recognized browsers. */
const ENGINES: Readonly<Record<string, string>> = {
Firefox: "Gecko",
LibreWolf: "Gecko",
Safari: "Safari",
}
/** Oldest plausible major versions (≈2023 releases); see the Python source. */
const ANCIENT: Readonly<Record<string, number>> = {
Firefox: 108,
Chrome: 108,
Edge: 108,
Opera: 95,
}
/** True when a `Browser/major` claims an impossibly old version. */
export function spoofed(b: string): boolean {
const i = b.indexOf("/")
const name = i < 0 ? b : b.slice(0, i)
const ver = i < 0 ? "" : b.slice(i + 1)
const floor = ANCIENT[name]
return floor !== undefined && /^\d+$/.test(ver) && Number(ver) < floor
}
/** Engine for a `Browser/major` result; Chromium is the modern default. */
export function engine(b: string): string {
const name = b.split("/")[0]
return ENGINES[name] ?? (name ? "Chromium" : "")
}
/** Desktop OS name, or "" when not recognizable. */
export function os(ua: string): string {
if (ua.includes("Windows NT")) return "Windows"
if (ua.includes("Mac OS X")) return "macOS"
if (ua.includes("Linux") || ua.includes("X11")) return "Linux"
if (ua.includes("Windows")) return "Windows"
if (ua.includes("Darwin")) return "macOS"
return ""
}
/**
* Human-readable phone name for an Android model code.
*
* Returns the input unchanged when nothing is known about it (Pixel and
* most other brands already send readable names).
*/
export function modelName(model: string): string {
if (model.startsWith("SM-")) {
// Strip the region/carrier suffix: SM-S918B -> SM-S918, and the
// Chinese/HK variant's trailing zero: SM-S9370 -> SM-S937.
let code = model.replace(/[A-Z]{1,2}$/, "")
if (!(code in SAMSUNG) && code.endsWith("0")) code = code.slice(0, -1)
if (code in SAMSUNG) return SAMSUNG[code]
const series = SAMSUNG_SERIES[code.slice(0, 4)]
if (series) return series
}
return model
}
const CACHE_MAX = 1024
const cache = new Map<string, UA>()
/** Parse a User-Agent string into a compact {@link UA} record. */
export function uaparse(ua: string | null | undefined): UA {
if (!ua || !ua.trim() || ua === "-" || ua === "null") return EMPTY
const hit = cache.get(ua)
if (hit) {
// Refresh recency, mirroring Python's lru_cache.
cache.delete(ua)
cache.set(ua, hit)
return hit
}
const r = parse(ua)
if (cache.size >= CACHE_MAX) {
cache.delete(cache.keys().next().value as string)
}
cache.set(ua, r)
return r
}
function parse(ua: string): UA {
const [name, kind] = bot(ua)
if (name) {
// The browser/OS in crawler UAs is a disguise; the bot identity is
// the relevant information, so `engine` and `os` are left empty.
let pretty = PRETTY_OVERRIDE[name]
if (pretty === undefined) {
const label = LABELED.has(name) ? (KIND_LABEL[kind] ?? "") : ""
pretty = label ? `${name} (${label})` : name
}
return {
...EMPTY,
pretty,
kind,
url: url(ua),
provider: PROVIDER_OF[name] ?? "",
}
}
return parseClient(ua)
}
function parseClient(ua: string): UA {
const r = client(ua)
// Frozen ancient browser strings are scanners/scripts, not users: show
// the claimed browser, but mark it and drop the fake engine/os/kind.
if (r.kind === "browser" && spoofed(browser(ua))) {
return { ...EMPTY, pretty: `${r.pretty} (spoofed)` }
}
return r
}
function client(ua: string): UA {
// Non-browser HTTP clients ("python-requests/2.32.5", "curl/8.0",
// "pip/24.3.1 {json…}"): the first product token, plus the OS when
// their payload mentions one in free text.
if (!ua.startsWith("Mozilla")) {
const token = ua.split(" ")[0]
let pretty = token.includes("/") ? token : ua
const osName = os(ua)
if (osName && !pretty.includes(osName)) pretty = `${pretty} ${osName}`
return { ...EMPTY, pretty, os: osName }
}
// HarmonyOS carries an "Android" compatibility token, so it must be
// detected before Android.
if (
ua.includes("OpenHarmony") ||
ua.includes("HarmonyOS") ||
ua.includes("ArkWeb")
) {
const b = browser(ua)
return {
...EMPTY,
pretty: b ? `${b} HarmonyOS` : "HarmonyOS",
engine: "ArkWeb",
os: "HarmonyOS",
kind: "browser",
}
}
if (ua.includes("iPhone") || ua.includes("iPad")) {
const device = ua.includes("iPhone") ? "iPhone" : "iPad"
const m = /OS (\d+)/.exec(ua)
return {
...EMPTY,
pretty: m ? `${device} iOS ${m[1]}` : device,
engine: "Safari",
os: "iOS",
kind: "browser",
}
}
const am = /Android ([\d.]+)/.exec(ua)
if (am) {
const b = browser(ua)
const mm = ANDROID_MODEL.exec(ua)
const token = mm ? mm[1].trim() : ""
let parts: string[]
if (token === "K") {
// Chrome's reduced UA freezes both: "Android 10; K". Neither
// is real — report just the OS.
parts = [b, "Android"].filter(Boolean)
} else {
// Firefox sends the form factor ("Mobile"/"Tablet") in the model
// slot. A known model replaces the OS.
let model = ""
if (token && !["wv", "Mobile", "Tablet"].includes(token)) {
model = modelName(token)
}
parts = [b, model || `Android ${am[1]}`].filter(Boolean)
}
return {
...EMPTY,
pretty: parts.join(" "),
engine: engine(b),
os: "Android",
kind: "browser",
}
}
const b = browser(ua)
const osName = os(ua)
const pretty = `${b} ${osName}`.trim()
return {
...EMPTY,
pretty: pretty || ua,
engine: engine(b),
os: osName,
kind: "browser",
}
}
+231
View File
@@ -0,0 +1,231 @@
// Generated by scripts/port_tables.py from uarite/bots.py and
// uarite/clients.py. Do not edit by hand; re-run the script.
/** UA token -> display name, checked in order; first hit wins. */
export const BROWSERS: ReadonlyArray<readonly [string, string]> = [
["HuaweiBrowser", "HuaweiBrowser"],
["EdgA", "Edge"],
["Edg", "Edge"],
["OPR", "Opera"],
["Vivaldi", "Vivaldi"],
["YaBrowser", "Yandex"],
["Brave", "Brave"],
["Whale", "Whale"],
["SamsungBrowser", "Samsung Internet"],
["MiuiBrowser", "Mi Browser"],
["UCBrowser", "UC Browser"],
["QQBrowser", "QQ Browser"],
["DuckDuckGo", "DuckDuckGo"],
["LibreWolf", "LibreWolf"],
["Firefox", "Firefox"],
["Chrome", "Chrome"],
["CriOS", "Chrome"],
["FxiOS", "Firefox"],
]
/** Lowercase UA substring -> [display name, kind]. First match wins. */
export const BOTS: Readonly<Record<string, readonly [string, string]>> = {
"mozilla/5.0 (x11; linux x86_64; rv:45.0) gecko/20100101 firefox/45.0": [
"Qualys SSL Labs",
"spider",
],
"feedfetcher-google": ["Feedfetcher-Google", "search"],
"google-inspectiontool": ["Google-InspectionTool", "search"],
"google-read-aloud": ["Google-Read-Aloud", "ai"],
"mediapartners-google": ["Mediapartners-Google", "analytics"],
"adsbot-google": ["AdsBot-Google", "analytics"],
"apis-google": ["APIs-Google", "spider"],
"storebot-google": ["Storebot-Google", "search"],
"google-extended": ["Google-Extended", "ai"],
googlebot: ["Googlebot", "search"],
googleother: ["GoogleOther", "ai"],
bingbot: ["Bingbot", "search"],
applebot: ["Applebot", "search"],
gptbot: ["GPTBot", "ai"],
"oai-searchbot": ["OAI-SearchBot", "search"],
"chatgpt-user": ["ChatGPT-User", "ai"],
"claude-searchbot": ["Claude-SearchBot", "search"],
claudebot: ["ClaudeBot", "ai"],
"claude-user": ["Claude-User", "ai"],
"perplexity-user": ["Perplexity-User", "ai"],
perplexitybot: ["PerplexityBot", "search"],
grokbot: ["GrokBot", "ai"],
bytespider: ["Bytespider", "ai"],
reflectionbot: ["Reflectionbot", "ai"],
"amzn-searchbot": ["Amzn-SearchBot", "search"],
amazonbot: ["Amazonbot", "search"],
ahrefsbot: ["AhrefsBot", "search"],
mj12bot: ["MJ12bot", "analytics"],
facebookexternalhit: ["Facebook", "social"],
"meta-externalagent": ["Meta-ExternalAgent", "ai"],
"meta-externalfetcher": ["Meta-ExternalFetcher", "ai"],
"meta-webindexer": ["Meta-WebIndexer", "search"],
bingpreview: ["BingPreview", "search"],
pinterest: ["Pinterest", "social"],
embedly: ["Embedly", "social"],
iframely: ["Iframely", "social"],
discordbot: ["Discord", "social"],
slackbot: ["Slack", "social"],
telegrambot: ["Telegram", "social"],
twitterbot: ["Twitter", "social"],
linkedinbot: ["LinkedIn", "social"],
whatsapp: ["WhatsApp", "social"],
headlesschrome: ["HeadlessChrome", "spider"],
uptimerobot: ["UptimeRobot", "analytics"],
pingdom: ["Pingdom", "analytics"],
}
/** Pretty suffixes for the kinds more precise than a generic spider. */
export const KIND_LABEL: Readonly<Record<string, string>> = {
ai: "AI",
search: "search",
social: "social",
analytics: "analytics",
}
/** Bot display name -> provider. */
export const PROVIDER_OF: Readonly<Record<string, string>> = {
GoogleOther: "Google",
"Mediapartners-Google": "Google",
Googlebot: "Google",
"Storebot-Google": "Google",
"APIs-Google": "Google",
"Google-Extended": "Google",
"Google-Read-Aloud": "Google",
"Feedfetcher-Google": "Google",
"AdsBot-Google": "Google",
"Google-InspectionTool": "Google",
"Claude-User": "Anthropic",
"Claude-SearchBot": "Anthropic",
ClaudeBot: "Anthropic",
GPTBot: "OpenAI",
"ChatGPT-User": "OpenAI",
"OAI-SearchBot": "OpenAI",
"Perplexity-User": "Perplexity",
PerplexityBot: "Perplexity",
Amazonbot: "Amazon",
"Amzn-SearchBot": "Amazon",
Bingbot: "Microsoft",
BingPreview: "Microsoft",
"Meta-ExternalFetcher": "Meta",
"Meta-WebIndexer": "Meta",
Facebook: "Meta",
"Meta-ExternalAgent": "Meta",
}
/** Per-bot pretty overrides: the full display string. */
export const PRETTY_OVERRIDE: Readonly<Record<string, string>> = {
Facebook: "Facebook",
"Feedfetcher-Google": "Google Feedfetcher (search)",
"Google-InspectionTool": "Google InspectionTool (search)",
"Google-Read-Aloud": "Google Read-Aloud (AI)",
"Mediapartners-Google": "Google Mediapartners (analytics)",
"AdsBot-Google": "Google AdsBot (analytics)",
"APIs-Google": "Google APIs",
"Storebot-Google": "Google Storebot (search)",
"Google-Extended": "Google Extended (AI)",
GoogleOther: "Google Other (AI)",
}
/** Bot names whose kind label is displayed. */
export const LABELED: ReadonlySet<string> = new Set([
"APIs-Google",
"AdsBot-Google",
"ChatGPT-User",
"Claude-SearchBot",
"Claude-User",
"ClaudeBot",
"Facebook",
"Feedfetcher-Google",
"GPTBot",
"Google-Extended",
"Google-InspectionTool",
"Google-Read-Aloud",
"GoogleOther",
"Googlebot",
"Mediapartners-Google",
"Meta-ExternalAgent",
"Meta-ExternalFetcher",
"Meta-WebIndexer",
"OAI-SearchBot",
"Perplexity-User",
"PerplexityBot",
"Storebot-Google",
])
/** Samsung model code (without region/carrier letter) -> marketing name. */
export const SAMSUNG: Readonly<Record<string, string>> = {
"SM-G930": "Galaxy S7",
"SM-G935": "Galaxy S7 Edge",
"SM-G950": "Galaxy S8",
"SM-G955": "Galaxy S8+",
"SM-G960": "Galaxy S9",
"SM-G965": "Galaxy S9+",
"SM-G970": "Galaxy S10e",
"SM-G973": "Galaxy S10",
"SM-G975": "Galaxy S10+",
"SM-G977": "Galaxy S10 5G",
"SM-G980": "Galaxy S20",
"SM-G981": "Galaxy S20",
"SM-G985": "Galaxy S20+",
"SM-G986": "Galaxy S20+",
"SM-G988": "Galaxy S20 Ultra",
"SM-G990": "Galaxy S21 FE",
"SM-G991": "Galaxy S21",
"SM-G996": "Galaxy S21+",
"SM-G998": "Galaxy S21 Ultra",
"SM-S901": "Galaxy S22",
"SM-S906": "Galaxy S22+",
"SM-S908": "Galaxy S22 Ultra",
"SM-S911": "Galaxy S23",
"SM-S916": "Galaxy S23+",
"SM-S918": "Galaxy S23 Ultra",
"SM-S921": "Galaxy S24",
"SM-S926": "Galaxy S24+",
"SM-S928": "Galaxy S24 Ultra",
"SM-S931": "Galaxy S25",
"SM-S936": "Galaxy S25+",
"SM-S937": "Galaxy S25 Edge",
"SM-S938": "Galaxy S25 Ultra",
"SM-S942": "Galaxy S26",
"SM-S946": "Galaxy S26+",
"SM-S948": "Galaxy S26 Ultra",
"SM-N930": "Galaxy Note 7",
"SM-N950": "Galaxy Note 8",
"SM-N960": "Galaxy Note 9",
"SM-N970": "Galaxy Note 10",
"SM-N975": "Galaxy Note 10+",
"SM-N980": "Galaxy Note 20",
"SM-N981": "Galaxy Note 20",
"SM-N985": "Galaxy Note 20 Ultra",
"SM-N986": "Galaxy Note 20 Ultra",
"SM-F700": "Galaxy Z Flip",
"SM-F707": "Galaxy Z Flip 5G",
"SM-F711": "Galaxy Z Flip3",
"SM-F721": "Galaxy Z Flip4",
"SM-F731": "Galaxy Z Flip5",
"SM-F741": "Galaxy Z Flip6",
"SM-F766": "Galaxy Z Flip7",
"SM-F900": "Galaxy Fold",
"SM-F907": "Galaxy Fold 5G",
"SM-F916": "Galaxy Z Fold2",
"SM-F926": "Galaxy Z Fold3",
"SM-F936": "Galaxy Z Fold4",
"SM-F946": "Galaxy Z Fold5",
"SM-F956": "Galaxy Z Fold6",
"SM-F966": "Galaxy Z Fold7",
}
/** Samsung series for codes missing from SAMSUNG. */
export const SAMSUNG_SERIES: Readonly<Record<string, string>> = {
"SM-S": "Galaxy S",
"SM-G": "Galaxy S",
"SM-N": "Galaxy Note",
"SM-A": "Galaxy A",
"SM-J": "Galaxy J",
"SM-M": "Galaxy M",
"SM-E": "Galaxy E",
"SM-F": "Galaxy Z",
"SM-T": "Galaxy Tab",
"SM-X": "Galaxy Tab",
}
+145
View File
@@ -0,0 +1,145 @@
import assert from "node:assert/strict"
import { test } from "node:test"
import { uaparse } from "../dist/index.js"
test("empty and missing UAs", () => {
for (const ua of ["", " ", "-", "null", null, undefined]) {
assert.deepEqual(uaparse(ua), {
pretty: "",
engine: "",
os: "",
kind: "",
url: "",
provider: "",
})
}
})
test("Chrome on Windows", () => {
const r = uaparse(
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36",
)
assert.equal(r.pretty, "Chrome/152 Windows")
assert.equal(r.engine, "Chromium")
assert.equal(r.os, "Windows")
assert.equal(r.kind, "browser")
assert.equal(r.url, "")
})
test("Safari on macOS", () => {
const r = uaparse(
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/18.1 Safari/605.1.15",
)
assert.equal(r.pretty, "Safari/18 macOS")
assert.equal(r.engine, "Safari")
})
test("Firefox on Android keeps the OS version", () => {
const r = uaparse(
"Mozilla/5.0 (Android 15; Mobile; rv:154.0) Gecko/154.0 Firefox/154.0",
)
assert.equal(r.pretty, "Firefox/154 Android 15")
assert.equal(r.engine, "Gecko")
assert.equal(r.os, "Android")
})
test("Chrome on Android with a Pixel model", () => {
const r = uaparse(
"Mozilla/5.0 (Linux; Android 12; Pixel 6) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Mobile Safari/537.36",
)
assert.equal(r.pretty, "Chrome/118 Pixel 6")
assert.equal(r.os, "Android")
})
test("Samsung model codes resolve to marketing names", () => {
const r = uaparse(
"Mozilla/5.0 (Linux; Android 13; SM-S918B) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Mobile Safari/537.36",
)
assert.equal(r.pretty, "Chrome/118 Galaxy S23 Ultra")
})
test("Chrome reduced UA (Android 10; K)", () => {
const r = uaparse(
"Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Mobile Safari/537.36",
)
assert.equal(r.pretty, "Chrome/152 Android")
})
test("iPhone", () => {
const r = uaparse(
"Mozilla/5.0 (iPhone; CPU iPhone OS 17_0 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Mobile/15E148 Safari/604.1",
)
assert.equal(r.pretty, "iPhone iOS 17")
assert.equal(r.engine, "Safari")
assert.equal(r.os, "iOS")
})
test("HarmonyOS before Android", () => {
const r = uaparse(
"Mozilla/5.0 (Linux; Android 12; HarmonyOS; ALN-AL00; HMSCore 6.13.0.312) AppleWebKit/537.36 (KHTML, like Gecko) HuaweiBrowser/6.0.1.311 Mobile Safari/537.36",
)
assert.equal(r.pretty, "HuaweiBrowser/6 HarmonyOS")
assert.equal(r.engine, "ArkWeb")
assert.equal(r.os, "HarmonyOS")
})
test("ancient browser versions are marked spoofed", () => {
const r = uaparse(
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/60.0.0.0 Safari/537.36",
)
assert.equal(r.pretty, "Chrome/60 Windows (spoofed)")
assert.equal(r.kind, "")
})
test("HTTP clients", () => {
assert.equal(
uaparse("python-requests/2.32.5").pretty,
"python-requests/2.32.5",
)
assert.equal(uaparse("curl/8.0").pretty, "curl/8.0")
})
test("GPTBot", () => {
const r = uaparse(
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot",
)
assert.equal(r.pretty, "GPTBot (AI)")
assert.equal(r.kind, "ai")
assert.equal(r.provider, "OpenAI")
assert.equal(r.url, "https://openai.com/gptbot")
})
test("Facebook external hit, disguised as Chrome on Android", () => {
const r = uaparse(
"Mozilla/5.0 (Linux; Android 13; Pixel 7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/134.0.6885.65 Mobile Safari/537.36; compatible; facebookexternalhit/1.1; +http://www.facebook.com/externalhit_uatext.php",
)
assert.equal(r.pretty, "Facebook")
assert.equal(r.kind, "social")
assert.equal(r.provider, "Meta")
assert.equal(r.url, "http://www.facebook.com/externalhit_uatext.php")
})
test("unknown bot from product token", () => {
const r = uaparse("NewBot/1.0 (+https://example.com/bot)")
assert.equal(r.pretty, "NewBot")
assert.equal(r.kind, "spider")
})
test("named from compatible token", () => {
const r = uaparse(
"Mozilla/5.0 (compatible; Page2RSS/0.7; +http://page2rss.com/)",
)
assert.equal(r.pretty, "Page2RSS")
assert.equal(r.kind, "spider")
})
test("unrecognized Mozilla UA falls back to the raw string", () => {
const ua = "Mozilla/5.0 (something entirely unknown)"
assert.equal(uaparse(ua).pretty, ua)
})
test("results are cached", () => {
const ua =
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36"
assert.equal(uaparse(ua), uaparse(ua))
})
+14
View File
@@ -0,0 +1,14 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "NodeNext",
"moduleResolution": "NodeNext",
"lib": ["ES2022"],
"outDir": "dist",
"declaration": true,
"strict": true,
"verbatimModuleSyntax": true,
"skipLibCheck": true
},
"include": ["src"]
}
+80 -42
View File
@@ -1,19 +1,24 @@
"""Crawler and link-preview tables."""
"""Crawler and link-unfurler tables."""
#: Lowercase UA substring to (display name, kind). Ordered: first match
#: wins, so overlapping names go from most to least specific.
BOTS = {
# Qualys SSL Labs scanner: frozen on this exact Firefox/45 string since
# ~2016; the pinned Gecko date makes the substring distinctive.
"mozilla/5.0 (x11; linux x86_64; rv:45.0) gecko/20100101 firefox/45.0": (
"Qualys SSL Labs",
"spider",
),
"feedfetcher-google": ("Feedfetcher-Google", "search"),
"google-inspectiontool": ("Google-InspectionTool", "search"),
"google-read-aloud": ("Google-Read-Aloud", "search"),
"mediapartners-google": ("Mediapartners-Google", "spider"),
"adsbot-google": ("AdsBot-Google", "spider"),
"google-read-aloud": ("Google-Read-Aloud", "ai"),
"mediapartners-google": ("Mediapartners-Google", "analytics"),
"adsbot-google": ("AdsBot-Google", "analytics"),
"apis-google": ("APIs-Google", "spider"),
"storebot-google": ("Storebot-Google", "spider"),
"duplexweb-google": ("DuplexWeb-Google", "spider"),
"storebot-google": ("Storebot-Google", "search"),
"google-extended": ("Google-Extended", "ai"),
"googlebot": ("Googlebot", "search"),
"googleother": ("GoogleOther", "spider"),
"googleother": ("GoogleOther", "ai"),
"bingbot": ("Bingbot", "search"),
"applebot": ("Applebot", "search"),
"gptbot": ("GPTBot", "ai"),
@@ -29,34 +34,39 @@ BOTS = {
"reflectionbot": ("Reflectionbot", "ai"),
"amzn-searchbot": ("Amzn-SearchBot", "search"),
"amazonbot": ("Amazonbot", "search"),
"ahrefsbot": ("AhrefsBot", "spider"),
"mj12bot": ("MJ12bot", "spider"),
"facebookexternalhit": ("Facebook", "preview"),
"meta-externalagent": ("Meta", "preview"),
"skypeuripreview": ("Skype", "preview"),
"bingpreview": ("BingPreview", "preview"),
"pinterest": ("Pinterest", "preview"),
"embedly": ("Embedly", "preview"),
"iframely": ("Iframely", "preview"),
"discordbot": ("Discord", "preview"),
"slackbot": ("Slack", "preview"),
"telegrambot": ("Telegram", "preview"),
"twitterbot": ("Twitter", "preview"),
"linkedinbot": ("LinkedIn", "preview"),
"whatsapp": ("WhatsApp", "preview"),
"ahrefsbot": ("AhrefsBot", "search"),
"mj12bot": ("MJ12bot", "analytics"),
"facebookexternalhit": ("Facebook", "social"),
"meta-externalagent": ("Meta-ExternalAgent", "ai"),
"meta-externalfetcher": ("Meta-ExternalFetcher", "ai"),
"meta-webindexer": ("Meta-WebIndexer", "search"),
"bingpreview": ("BingPreview", "search"),
"pinterest": ("Pinterest", "social"),
"embedly": ("Embedly", "social"),
"iframely": ("Iframely", "social"),
"discordbot": ("Discord", "social"),
"slackbot": ("Slack", "social"),
"telegrambot": ("Telegram", "social"),
"twitterbot": ("Twitter", "social"),
"linkedinbot": ("LinkedIn", "social"),
"whatsapp": ("WhatsApp", "social"),
"headlesschrome": ("HeadlessChrome", "spider"),
"phantomjs": ("PhantomJS", "spider"),
"uptimerobot": ("UptimeRobot", "spider"),
"pingdom": ("Pingdom", "spider"),
"uptimerobot": ("UptimeRobot", "analytics"),
"pingdom": ("Pingdom", "analytics"),
}
#: Pretty suffixes for the kinds more precise than a generic spider.
KIND_LABEL = {"ai": "AI", "search": "search", "preview": "preview"}
KIND_LABEL = {
"ai": "AI",
"search": "search",
"social": "social",
"analytics": "analytics",
}
#: Providers with more than one crawler product; the kind label is kept
#: only where it distinguishes siblings within the group.
PROVIDERS = [
(
#: Crawler product families: provider -> the display names of its bots.
#: The kind label is kept only where it distinguishes siblings within a family.
PROVIDERS = {
"Google": frozenset({
"Googlebot",
"Google-Extended",
"GoogleOther",
@@ -67,19 +77,47 @@ PROVIDERS = [
"AdsBot-Google",
"APIs-Google",
"Storebot-Google",
"DuplexWeb-Google",
),
("ClaudeBot", "Claude-User", "Claude-SearchBot"),
("GPTBot", "OAI-SearchBot", "ChatGPT-User"),
("PerplexityBot", "Perplexity-User"),
("Amazonbot", "Amzn-SearchBot"),
("Bingbot", "BingPreview"),
]
}),
"Anthropic": frozenset({"ClaudeBot", "Claude-User", "Claude-SearchBot"}),
"OpenAI": frozenset({"GPTBot", "OAI-SearchBot", "ChatGPT-User"}),
"Perplexity": frozenset({"PerplexityBot", "Perplexity-User"}),
"Amazon": frozenset({"Amazonbot", "Amzn-SearchBot"}),
"Microsoft": frozenset({"Bingbot", "BingPreview"}),
"Meta": frozenset({
"Facebook",
"Meta-ExternalAgent",
"Meta-ExternalFetcher",
"Meta-WebIndexer",
}),
}
#: Bot names whose kind label is displayed, computed from the groups.
#: Reverse lookup: bot display name -> provider.
PROVIDER_OF = {
name: provider for provider, names in PROVIDERS.items() for name in names
}
#: Per-bot pretty overrides: the full display string, replacing the
#: name-plus-kind-label composition entirely.
PRETTY_OVERRIDE = {
"Facebook": "Facebook",
"Feedfetcher-Google": "Google Feedfetcher (search)",
"Google-InspectionTool": "Google InspectionTool (search)",
"Google-Read-Aloud": "Google Read-Aloud (AI)",
"Mediapartners-Google": "Google Mediapartners (analytics)",
"AdsBot-Google": "Google AdsBot (analytics)",
"APIs-Google": "Google APIs",
"Storebot-Google": "Google Storebot (search)",
"Google-Extended": "Google Extended (AI)",
"GoogleOther": "Google Other (AI)",
}
#: Reverse lookup: bot display name -> kind.
NAME_KIND = {name: kind for name, kind in BOTS.values()}
#: Bot names whose kind label is displayed: those in families with mixed kinds.
LABELED = {
name
for group in PROVIDERS
if len({kind for n, kind in BOTS.values() if n in group}) > 1
for name in group
for names in PROVIDERS.values()
if len({NAME_KIND[name] for name in names}) > 1
for name in names
}
+16 -10
View File
@@ -58,7 +58,11 @@ SAMSUNG = {
"SM-S928": "Galaxy S24 Ultra",
"SM-S931": "Galaxy S25",
"SM-S936": "Galaxy S25+",
"SM-S937": "Galaxy S25 Edge",
"SM-S938": "Galaxy S25 Ultra",
"SM-S942": "Galaxy S26",
"SM-S946": "Galaxy S26+",
"SM-S948": "Galaxy S26 Ultra",
# Galaxy Note
"SM-N930": "Galaxy Note 7",
"SM-N950": "Galaxy Note 8",
@@ -69,20 +73,22 @@ SAMSUNG = {
"SM-N981": "Galaxy Note 20",
"SM-N985": "Galaxy Note 20 Ultra",
"SM-N986": "Galaxy Note 20 Ultra",
# Galaxy Z foldables
# Galaxy Z foldables (Samsung dropped the space from Fold2/Flip3 on)
"SM-F700": "Galaxy Z Flip",
"SM-F707": "Galaxy Z Flip 5G",
"SM-F711": "Galaxy Z Flip 3",
"SM-F721": "Galaxy Z Flip 4",
"SM-F731": "Galaxy Z Flip 5",
"SM-F741": "Galaxy Z Flip 6",
"SM-F711": "Galaxy Z Flip3",
"SM-F721": "Galaxy Z Flip4",
"SM-F731": "Galaxy Z Flip5",
"SM-F741": "Galaxy Z Flip6",
"SM-F766": "Galaxy Z Flip7",
"SM-F900": "Galaxy Fold",
"SM-F907": "Galaxy Fold 5G",
"SM-F916": "Galaxy Z Fold 2",
"SM-F926": "Galaxy Z Fold 3",
"SM-F936": "Galaxy Z Fold 4",
"SM-F946": "Galaxy Z Fold 5",
"SM-F956": "Galaxy Z Fold 6",
"SM-F916": "Galaxy Z Fold2",
"SM-F926": "Galaxy Z Fold3",
"SM-F936": "Galaxy Z Fold4",
"SM-F946": "Galaxy Z Fold5",
"SM-F956": "Galaxy Z Fold6",
"SM-F966": "Galaxy Z Fold7",
}
#: Samsung series for codes missing from the table above.
+73 -29
View File
@@ -4,16 +4,18 @@ import re
from dataclasses import dataclass
from functools import lru_cache
from .bots import BOTS, KIND_LABEL, LABELED
from .bots import BOTS, KIND_LABEL, LABELED, PRETTY_OVERRIDE, PROVIDER_OF
from .clients import BROWSERS, SAMSUNG, SAMSUNG_SERIES
@dataclass(frozen=True)
class UA:
pretty: str
bot: str
kind: str
url: str
pretty: str = ""
engine: str = ""
os: str = ""
kind: str = ""
url: str = ""
provider: str = ""
#: Fallback for unknown crawlers: a product token whose name says so.
@@ -46,7 +48,7 @@ def url(ua: str) -> str:
def bot(ua: str) -> tuple[str, str]:
"""(display name, kind) of the crawler/previewer the UA claims."""
"""(display name, kind) of the crawler/unfurler the UA claims."""
low = ua.lower()
m = BOTS_RE.search(low)
if m:
@@ -93,6 +95,29 @@ def browser(ua: str) -> str:
return ""
#: Engines that differ from the Chromium default for recognized browsers.
ENGINES = {"Firefox": "Gecko", "LibreWolf": "Gecko", "Safari": "Safari"}
#: Oldest plausible major versions (≈2023 releases). These browsers
#: auto-update, so anything older is a scanner's frozen string, not a real
#: installation. Safari is OS-tied and exempt; old Macs genuinely run it.
#: Bump the floors every few years.
ANCIENT = {"Firefox": 108, "Chrome": 108, "Edge": 108, "Opera": 95}
def spoofed(b: str) -> bool:
"""True when a ``Browser/major`` claims an impossibly old version."""
name, _, ver = b.partition("/")
floor = ANCIENT.get(name)
return floor is not None and ver.isdigit() and int(ver) < floor
def engine(b: str) -> str:
"""Engine for a ``Browser/major`` result; Chromium is the modern default."""
name = b.split("/")[0]
return ENGINES.get(name, "Chromium" if name else "")
def os(ua: str) -> str:
"""Desktop OS name, or "" when not recognizable."""
if "Windows NT" in ua:
@@ -115,8 +140,11 @@ def model_name(model: str) -> str:
most other brands already send readable names).
"""
if model.startswith("SM-"):
# Strip the trailing region/carrier letter: SM-S918B -> SM-S918.
# Strip the region/carrier suffix: SM-S918B -> SM-S918, and the
# Chinese/HK variant's trailing zero: SM-S9370 -> SM-S937.
code = re.sub(r"[A-Z]{1,2}$", "", model)
if code not in SAMSUNG and code.endswith("0"):
code = code[:-1]
if code in SAMSUNG:
return SAMSUNG[code]
series = SAMSUNG_SERIES.get(code[:4])
@@ -125,31 +153,45 @@ def model_name(model: str) -> str:
return model
def uaparse(ua: str) -> UA:
@lru_cache(maxsize=1024)
def uaparse(ua: str | None) -> UA:
"""Parse a User-Agent string into a compact ``UA`` record.
``pretty`` is "" for empty/missing UAs and the original string when
nothing is recognized.
Only the browser path is cached: real visitors repeat (cache hits),
while crawlers and scripts are mostly one-hit wonders whose entries
would just flush the cache.
Everything is cached: a dict lookup on the full UA string is far
cheaper than re-parsing, and the frozen ``UA`` is shared safely.
"""
if not ua or not ua.strip() or ua in ("-", "null"):
return UA("", "", "", "")
return UA()
name, kind = bot(ua)
if name:
# The browser/OS in crawler UAs is a disguise; the bot identity is
# the relevant information.
label = KIND_LABEL.get(kind, "") if name in LABELED else ""
pretty = f"{name} ({label})" if label else name
return UA(pretty, name, kind, url(ua))
# the relevant information, so ``engine`` and ``os`` are left empty.
pretty = PRETTY_OVERRIDE.get(name)
if pretty is None:
label = KIND_LABEL.get(kind, "") if name in LABELED else ""
pretty = f"{name} ({label})" if label else name
return UA(
pretty=pretty, kind=kind, url=url(ua),
provider=PROVIDER_OF.get(name, ""),
)
return _parse_client(ua)
@lru_cache(maxsize=1024)
def _parse_client(ua: str) -> UA:
"""Browser/client parsing behind the cache; ``uaparse`` filters bots out."""
"""Browser/client parsing; ``uaparse`` filters bots out."""
r = _client(ua)
# Frozen ancient browser strings are scanners/scripts, not users: show
# the claimed browser, but mark it and drop the fake engine/os/kind.
if r.kind == "browser" and spoofed(browser(ua)):
return UA(pretty=f"{r.pretty} (spoofed)")
return r
def _client(ua: str) -> UA:
"""Browser/client parsing; spoof marking is done by the caller."""
# Non-browser HTTP clients ("python-requests/2.32.5", "curl/8.0",
# "pip/24.3.1 {json…}"): the first product token, plus the OS when
@@ -160,19 +202,22 @@ def _parse_client(ua: str) -> UA:
os_name = os(ua)
if os_name and os_name not in pretty:
pretty = f"{pretty} {os_name}"
return UA(pretty, "", "", "")
return UA(pretty=pretty, os=os_name)
# HarmonyOS carries an "Android" compatibility token, so it must be
# detected before Android.
if "OpenHarmony" in ua or "HarmonyOS" in ua or "ArkWeb" in ua:
b = browser(ua)
return UA(f"{b} HarmonyOS" if b else "HarmonyOS", "", "browser", "")
return UA(
pretty=f"{b} HarmonyOS" if b else "HarmonyOS",
engine="ArkWeb", os="HarmonyOS", kind="browser",
)
if "iPhone" in ua or "iPad" in ua:
device = "iPhone" if "iPhone" in ua else "iPad"
m = re.search(r"OS (\d+)", ua)
pretty = f"{device} iOS {m.group(1)}" if m else device
return UA(pretty, "", "browser", "")
return UA(pretty=pretty, engine="Safari", os="iOS", kind="browser")
m = re.search(r"Android ([\d.]+)", ua)
if m:
@@ -190,13 +235,12 @@ def _parse_client(ua: str) -> UA:
if token and token not in ("wv", "Mobile", "Tablet"):
model = model_name(token)
parts = [p for p in (b, model or f"Android {m.group(1)}") if p]
return UA(" ".join(parts), "", "browser", "")
return UA(
pretty=" ".join(parts), engine=engine(b), os="Android",
kind="browser",
)
b = browser(ua)
pretty = f"{b} {os(ua)}".strip()
return UA(pretty or ua, "", "browser", "")
def is_bot(ua: str) -> bool:
"""True when the UA claims a crawler or link-preview identity."""
return bool(uaparse(ua).bot)
os_name = os(ua)
pretty = f"{b} {os_name}".strip()
return UA(pretty=pretty or ua, engine=engine(b), os=os_name, kind="browser")