corpus.blog/about

About corpus.blog

The citation index for the independent developer blogosphere: what independent developers publish, who points at it, and why every figure counts distinct blogs rather than articles.

corpus.blog is the citation index for the independent developer blogosphere. It records what independent developers publish, and who points at it.

Every figure on this site counts distinct blogs, never articles. The top one percent of sources produce a third of everything we hold, so ranking by volume would just return whichever site posts most, forever. Counting each blog once asks the question a citation was always meant to answer: how many different people found this worth pointing at. The full method is public and reproducible.

Two kinds of page make up the public record: one for each post we hold, and one for each blog. Both show who cites what, when, and with which words. What they do not show is the text. The text of a post is a private index we use to compute these figures, and it is never republished. A library's catalogue is public; its stacks are not. Read the post itself at its own address, or at its Wayback Machine capture.

The crawler fetches your feed, and one page per post your feed names, at a slow and declared pace. A blog with no feed is not fetched at all. It names itself so you can tell it apart from a browser, and block it if you want.

Claiming a blog is proof of control, not payment: one tag on your homepage, and the record of your blog becomes yours to correct and yours to draw from — who cites you, webmentions sent on your behalf, your own archive back as Markdown. Nothing about it costs anything, now or later.

A blog that is not on here can be added by anyone, with no account: one domain, and it joins the list of domains waiting to be read. Every other way a blog is found here is one we run ourselves, and the people who know about a blog nobody has listed yet are its readers and its author.

corpus.blog also keeps its own blog, with a feed like everyone else's.