Skip to content
98.39%
Top-1 accuracy
99.46%
Top-3 accuracy
0.035
Calibration error
186
Labeled domains

Top-1 means our highest-confidence answer is the correct operator, on 183 of 186 domains. Top-3 means the correct operator is somewhere in our top three. These are measured on a curated, labeled set, so treat them as an upper bound. Real-world accuracy varies with how hard a domain is, which is exactly what the table below is for.

Read this first

A benchmark is not the real world

The number that matters is not the headline, it is how the number behaves when a domain fights back. A clean corporate site with matching records is easy. A GDPR-redacted domain behind a shared CDN whose page still shows a pre-acquisition brand is not. We split the benchmark into difficulty categories precisely so a single average cannot hide the hard cases. When you run your own domains, expect results closer to the harder rows than the easy ones.

By difficulty

Where it holds, and where it slips

Sorted hardest first. The easy categories sit at 100%, which is expected and not very interesting. The rows worth reading are the ones below 100%, and how large a share of the set they represent.

Category Domains Top-1 Top-3
Known-hard
Cases we track precisely because they are difficult, including ones we still miss.
14 85.7% 92.9%
Subsidiary
A brand operated by a parent that markets under a different name.
10 90.0% 100.0%
Clear
WHOIS, certificate and hosting all agree on one operator.
56 100.0% 100.0%
Privacy-redacted
GDPR or proxy has stripped the registration record. Resolved from other signals.
17 100.0% 100.0%
CDN-masked
The domain sits behind a shared CDN that hides the origin.
15 100.0% 100.0%
Ambiguous
Several valid answers exist (brand, parent, subsidiary).
13 100.0% 100.0%
Acquisition
A brand acquired by another company, where the page may still show the old name.
11 100.0% 100.0%
CDN + conflict
CDN infrastructure points one way, ownership signals another.
10 100.0% 100.0%
Privacy + conflict
Redacted record plus signals that disagree with each other.
10 100.0% 100.0%
Rebrand
A company that changed names, where records and content can lag.
10 100.0% 100.0%
Signal conflict
Independent sources name different entities.
10 100.0% 100.0%
Single signal
Only one usable source is available for the whole domain.
10 100.0% 100.0%

The set has grown over time and now leans toward easier categories, which pulls the headline average up. That is why the honest signal is the per-category row, not the single number. Amber marks any category we do not yet solve perfectly.

Calibration

When we say 90%, is it right 90% of the time?

A confidence score is only useful if it means something. We bucket every prediction by the confidence we reported, then check how often that bucket was actually correct. The gap between the two, averaged across buckets, is the expected calibration error: 0.0354, or about 3.5 points.

Confidence range Predictions We said Actually right
24.0% to 90.0% 38 78.5% 94.7%
90.0% to 100.0% 37 96.2% 97.3%
100.0% to 100.0% 37 100.0% 100.0%
100.0% to 100.0% 37 100.0% 100.0%
100.0% to 100.0% 37 100.0% 100.0%

The lowest-confidence bucket is deliberately cautious: we reported an average 78.5% there and were actually right 94.7% of the time. We would rather round down and beat the score than the reverse, because overconfidence is the failure mode that burns an investigation.

Where we're wrong

3 misses out of 186

Every current top-1 miss, in full. Most are genuinely contested: a subsidiary attributed to its parent, or a name that shifted after an acquisition. We keep them in the set on purpose, because a benchmark you can pass by deleting the hard cases is worthless.

mullvad.net known regression
We returned
mullvad vpn
Correct answer
mullvad vpn ab
mandiant.com subsidiary mismatch
We returned
google
Correct answer
mandiant
espn.com known regression
We returned
disney walt disney company
Correct answer
espn
Methodology

How the number is made

The short version is below. For the full protocol, the exact match function, the labeling rules, the baseline, and the in-sample vs real-world distinction, read the complete methodology.

The set

186 domains hand-labeled across 12 difficulty categories, from clean corporate sites to redacted, CDN-masked, and post-acquisition edge cases. Each label is tied to an independent source, a corporate website, an SEC filing, a certificate transparency log, or a registration record, with the date it was verified.

The measure

For each domain we score every signal, rank the candidate operators, and check whether the correct one is our first answer (top-1) and whether it lands in our first three (top-3). Calibration error compares reported confidence against how often that confidence was right.

The cadence

The benchmark reruns on every change to the scoring engine, so a regression cannot ship quietly. The run is deterministic: the same inputs always produce the same score, which is what makes these numbers auditable rather than a lucky sample.

What we hold back

We do not publish the raw domain list, because it is our regression set and publishing it invites gaming. We also do not publish per-signal weights or the vendors behind each source. The outcomes are open; the recipe stays ours.

Run it on a domain you already know

The fastest way to trust a benchmark is to break it. Pick a domain whose operator you know cold, and see what comes back.