v2.3
en.wikipedia · disambiguation pageview analysis
Type an ambiguous term and see how Wikipedia’s pageviews divide among the topics that share it, over any run of months since 2015.
Try:
—
—
Articles
—
Views · 12 mo
Top share
—
Behaves like
| # | Article | Views · 12 mo | Share |
|---|---|---|---|
| Awaiting first analysis | |||
These pages collect many more articles sharing the term. Analyzing one runs the same pipeline on that list.
Listed per MOS:DABMENTION, because each topic is mentioned in another article and its entry links there.
A disambiguation page lists every topic with a claim to a name, one line each, in no particular order. Reader attention is distributed very differently. Territory measures how Wikipedia pageviews divide among the topics on that list, over any window since July 2015. The result is usually lopsided, and fifty listed topics often behave like a contest between three or four, which is why the readout gives both counts.
Wikipedia’s rule for who gets the bare name, WP:PRIMARYTOPIC, turns on two things. The first is usage, meaning whether one topic is much more likely than all the others combined. The second is long-term significance, which is a judgment call and can justify a title the topic does not dominate by traffic. Only usage is measurable, and it is the only one of the two this tool touches.
So when the article holding a name is not the one readers go to, Territory reports the disagreement and stops. What the title should be is a question for a Requested Move discussion, where usage is one input among several.
An article’s full readership stands in for attention to the name. Freddie Mercury’s readers count toward “Mercury” whether the name brought them there or not, which overstates any topic famous for reasons unrelated to what it is called. That is the unit this tool works in, and it is the largest assumption on the page.
A sharper instrument exists. Wikimedia publishes clickstream data showing which articles readers navigate to from a given disambiguation page, which would measure name-driven attention directly. It ships as a monthly bulk dump rather than an API, and Territory is a static page with no backend, so using it would mean building infrastructure this version does not have.
Each entry’s subject link, which MOS:DAB formatting puts first on the line, becomes a candidate article. Three kinds of entry are excluded, and every exclusion is counted somewhere you can see it.
Topics without their own article (MOS:DABMENTION). Prose before the first link means the subject is unlinked and the link only points at something the subject is mentioned in. Counting Warsaw’s readership toward “Phoenix”, because Warsaw is nicknamed Phoenix City, would measure something else entirely. These are listed under the results table.
Anchor entries, which link into a section of a larger article. Pageviews are recorded per page, so a section has no count of its own.
Container pages. This covers “List of…” pages, discographies, name pages, and set index articles, detected through Wikipedia’s own tracking category. They never occupy a row in the register, and appear instead as separate indexes you can run the tool on.
Every entry is resolved through redirects, fifty at a time, and the exclusion rules above run again on the destinations. An entry that forwards into a section, a list, or another disambiguation page is dropped, with the reason recorded in the browser console. Where several entries converge on one article they are merged, and the row keeps the page’s original wording under “listed as”.
Pageviews are counted against the title requested, so an article that receives substantial traffic through a redirect will undercount here. Summing redirect views is not implemented yet.
Counts come from the Wikimedia REST per-article API at monthly granularity, which is the finest resolution the API offers. One request per article fetches its whole recorded history, from July 2015 when the API begins through the last complete month. Partial months are excluded, because a partial total would corrupt every share it touches. Windows are cut from that series in the browser, so moving the window refetches nothing.
Bots are excluded and all platforms are combined. The upstream service intermittently returns “not found” for a rotating minority of titles, so each article runs a ladder before a zero is believed: human-only, a retry, a bot-inclusive fallback, then a second salvage pass over anything still missing. Any article whose count carries a caveat says so on its own row, whether that is included bot traffic, no data returned, or a fetch that failed outright. Failed fetches show no number and are left out of the share denominator. Shares are rounded to one decimal, so a column may not sum to exactly 100%.
The second number in the readout is how many equally-read articles would produce the same concentration of attention. It sits near 1 when one topic takes almost everything, and near the full count when attention spreads evenly. It is the inverse Herfindahl–Hirschman index of the shares in the table, computed over the articles with usable pageview data.
Every figure on the page is computed over the selected window and nothing else, including the totals, the shares, the concentration count, and the usage check. The window can be any run of months from July 2015 onward, down to a single month. Fig. 3 is the one exception. It always draws the full history with the selected window lit and the rest washed out, so the table answers who leads in the window while the chart shows where that window sits.
Short windows measure a news cycle rather than a settled readership, since a film’s release month will beat a century-old novel. Below three months the usage check says so on its own line. To find when the lead changed hands, move the window and re-read the table.
History is measured against today’s article list. The disambiguation page is read as it stands now, then priced backward through time. Topics that once shared the name but have since been renamed, merged, or delisted do not appear in earlier months, and an article younger than the chart begins when it begins. The early years of Fig. 3 show today’s claimants competing in yesterday’s traffic, which is a different thing from showing yesterday’s disambiguation page.
Miscounts, missing entries, and wrong verdicts are the useful reports. The exclusion rules are heuristics running against pages humans wrote by hand, so they will be wrong on some of them.
Thanks, that came through.
Worked examples: Georgia · Mercury · Rock. v2.3, September 2026.
< > [ ] { } and |
in page titles, and the API reports those as invalid rather than
missing. The resolver checked only for missing, so it treated the
invalid ones as real and reported articles that cannot exist. They
now get the same answer as a typo.