Report · Method

How we count

Every figure on this site can be explained in a sentence beside it, and every ranking can be reproduced from the rule below. A score nobody can reproduce is not shown.

1. We count blogs, not posts.

The top 1% of blogs produce a third of the corpus, so counting posts returns whoever posts most, forever. Each blog counts once per thing it points at.

2. A link from a page's own site is not a citation.

Self-citations are excluded. Contact pages, site chrome and navigation are excluded through the substantive-page rule.

3. Dates are the ones the blog claims.

The published date, never the day we stored it. Dates before 1990 and more than a day ahead are treated as missing, and a post with no date falls out of every windowed figure. How many that is stands in the holdings below.

4. The text is private.

The texts we hold are there to compute things; nothing shows or sells them. A record links to the post and its Wayback capture. An author can download their own.

5. Trends read English-language blogs only.

Language is a property of the blog, guessed from hundreds of titles. A blog that publishes changelogs cannot be classified and is not assumed English. A source is counted only once it published before the window opened, so a blog new to the corpus cannot read as a spike.

What a citation is, beyond rule 2

A link that is template furniture — a share button, a footer badge, the same link repeated on every post — is excluded the same way a self-citation is: a widget on every page is not an author's choice. Links to the infrastructure every blog links to — Wikipedia, GitHub, YouTube and the like — stay in the record, so a blog's own page still lists them, but they are excluded from the rankings: without that exclusion, the most-cited leaderboard is just a list of the same handful of platforms, every week, forever.

Talked about

Terms are pulled from post titles, and a blog counts once per term regardless of how many titles it used. A term is trending when its recent breadth — how many blogs used it in the last week — has grown against a twelve-week baseline, both figures normalised by how many blogs published at all in that window, so a quiet week does not read as every term collapsing at once. The ranking is breadth times lift: breadth alone returns the same vocabulary every week, lift alone crowns whatever five blogs just coined. A term needs at least five blogs behind it to count, and only blogs identified as English are read. Sources identified as something other than a blog — documentation, changelogs, news outlets, repositories, directories, podcasts — are excluded, and so is a longer-than-usual stopword list, including every month and weekday: they trend every single week by definition, and a calendar is not a story.

The query behind most cited

Most cited does not run from hand-written SQL kept here in sync with nothing — it runs from App\Actions\Citations\GetMostCitedArticles, the one place that query is allowed to live. In words: for each distinct cited page, count the distinct blogs whose posts link to it within the window — 30 days on the front page — excluding a blog citing its own page, excluding chrome and the infrastructure every blog links to, and excluding a bare homepage, which names a site rather than a piece of writing. A raw link count rides along beside the rank, since one blog linking somewhere fifty times is worth seeing next to the number that actually orders the list.

The same record is available as data, not just as a page — see the public API.

What we hold

Counted when this page was built, not typed in: the same figures the masthead above carries.

Blogs
0
Posts
0
Links between blogs we hold
0
Posted in the last 7 days
0
Posted in the last 30 days
0
Posts with no date
0