Proper-cased display forms of each source's internal (lowercase) POPULARITY_METRIC/repo_tbl$source key, for anywhere a source name is shown to a reader rather than matched against data (e.g. plot annotations). "npm" is genuinely lowercase as a name, not an abbreviation, so it's left as-is.
Fit a quasi-Poisson GLM to analyse differences in monthly issue rates across popularity strata. The month_num:popularity_stratum interaction tests whether the long tail trends differently from the popular head, rather than just reporting one global trend line.
Extract every issue (pull requests excluded) opened against a single GitHub repo, with the opener's handle and a contribution score: that handle's fractional share (0-1) of all commits ever landed on the repo's default branch, or 0 if the author isn't a contributor at all (see the note at the top of this file for why this is a coarser but far cheaper substitute for "was this author already a contributor at the time they opened the issue"). Also carries each issue's comment count, and the repo's own GitHub creation timestamp, used elsewhere as the start of a repo's exposure window.
Build the (popularity stratum-x-month) issue-rate table for one source: a chosen metric from non-contributor issues, aggregated over a trailing rolling window of months and normalized by repo-months of exposure, where a repo's exposure begins at its GitHub creation date. "Non-contributor" here means contribution <= contrib_threshold (see github_issue_authors() for how contribution - each author's fractional share of all commits ever landed on the repo - is computed). repo_created_at lives on issue_authors_tbl (fetched alongside each repo's issues by github_issue_authors()/fetch_issue_authors()), not repo_tbl, so a repo only contributes exposure once it's been fetched at least once - repos with zero issues fetched (either not yet fetched at all, or fetched and genuinely having none) don't have a repo_created_at on file and are excluded here rather than analysed.
Full npm monthly download-count population (~3.77M packages), via the download-counts npm package: https://www.npmjs.com/package/download-counts — a single static JSON object, republished monthly, mapping package name to last-month download count. This is npm's practical equivalent of PyPI's ClickHouse/BigQuery dataset: there is no direct npm counterpart of BigQuery's public PyPI download-log dataset.
Plot the trailing-window rate (issue_rate_tbl()'s rate column - see its metric param for whether that's issues or comments per repo-month) over time, one line per popularity stratum.
Compare monthly issue rate across all four sources (pypi, npm, joss, ropensci), for one popularity stratum. Note that "stratum" is relative to each source's own distribution (see issue_rate_tbl()/ popularity_strata()) - e.g. pypi's Q4 download count and joss's Q4 star count aren't the same absolute popularity, just each source's own top quarter. A source with no data yet for the requested window (e.g. not fully fetched - see analysis-plan.md) just contributes no line, rather than erroring.
Full PyPI download-count population (~870k packages, last complete calendar month), paginated in chunks of CLICKHOUSE_PAGE_SIZE. Typically ~9 requests, well under a minute, no rate limiting encountered.