Build a working sample from a full (name, downloads)
population, stratified by download popularity: a deterministic head of the
top_n_head most-downloaded packages, plus a random draw of up to
tail_size packages from the remaining long tail.
Usage
build_working_sample(
downloads_tbl = NULL,
top_n_head = 15000L,
tail_size = 40000L,
label = NULL
)Arguments
- downloads_tbl
A tibble with at least
nameanddownloadscolumns, as returned bypypi_downloads_full()ornpm_downloads_full().- top_n_head
Number of most-downloaded packages to include deterministically.
- tail_size
Number of packages to randomly draw from the remaining long tail (outside the head). Capped at the size of that tail.
- label
Optional string used to prefix a
clistatus message (e.g."PyPI"or"npm"); no message is printed ifNULL.