# Methodology

Every page on this site is generated from a structured corpus of public research metadata. This is how that corpus is built, and what its numbers do and do not mean.

## How records get here

- Records are ingested from public scholarly metadata APIs — principally OpenAlex. See [Sources](/sources/).
- Ingest is seeded from a set of publications, then expanded outward: the authors of those papers, the journals they appeared in, the institutions those authors are affiliated with, and the topics the papers belong to.
- Each entity keeps its upstream identifiers — OpenAlex id, and where present DOI, PMID, ORCID, ROR, ISSN, Wikidata. Those are on every page so a record can be traced back to its source rather than taken on trust.
- Links between records are real foreign keys in our database, not name matches. A paper is linked to a researcher because the source data links them, not because two strings looked similar.

## What is published, and what is held back

199,898 entity pages are published today. That is smaller than the corpus behind it: expansion is capped so the published set stays reviewable, and a record with no usable name is not given a page at all. Growth is deliberate rather than automatic — the published set only grows when the corpus behind it has been checked.

## Reading the numbers

- **Works** and **citations** are the upstream totals for that entity across all of the source database — not a count of what is published on this site.
- **h-index** and **i10-index** are taken from the source's own summary statistics. We do not recompute them.
- Lists on a page — top papers, collaborating institutions, publication-per-year charts, open-access breakdowns — are computed over the records shown on that page, which are capped. They describe that visible set, not the entity's whole output.
- A figure we do not have is left off the page. Blank means unknown, never zero.

## What this data cannot tell you

Citation counts measure attention, not quality, and they are unevenly distributed across fields and decades. Author-paper links are incomplete for very large collaborations, where source records truncate the author list. Affiliations reflect the last known institution in the source data, which lags reality. Nothing here should be used to rank or assess an individual researcher.

We are not a bibliometric authority and we do not publish a ranking. [The FAQ](/faq/) covers the questions this raises most often.
