Skip to contents

Unlike fetch_issue_authors(), which skips any repo already marked done, this re-fetches every repo in repo_urls, but each repo's own github_issue_authors() call is scoped with since set to that repo's most recent last_updated value already on disk - so GitHub only returns issues that are new, or that changed (e.g. picked up new comments) since that time. A repo not yet present in the checkpoint is fetched in full, exactly as fetch_issue_authors() would. Rows returned for an already-known issue replace the stale row; all other existing rows are left untouched.

Usage

update_issue_authors(repo_urls, out_dir, batch_size = 50L)

Arguments

repo_urls

Character vector of repo URLs to refresh. Repos not already present in the out_dir checkpoint are fetched in full.

out_dir

Directory holding the issue-authors.csv checkpoint written/read by fetch_issue_authors()/read_issue_authors_data().

batch_size

Repos refreshed (concurrently) per checkpoint write.

Value

A tibble with columns repo_url, issue_number, author, created_at, n_comments, contribution, repo_created_at, last_updated - the full accumulated result, with refreshed repos' rows brought up to date.

Details

Because each repo's since cursor advances every time it's refreshed, this is safe to re-run (e.g. from a scheduled job) without any separate "done" checkpoint: a run interrupted partway simply leaves the not-yet-reached repos with an older last_updated, picked up as normal on the next call. Batches are drawn via interlace_for_even_coverage() rather than sequentially, so an interrupted run leaves progress spread across repo_urls rather than concentrated at the top.

Examples

repo_urls <- c (
    "https://github.com/ropensci/targets",
    "https://github.com/ropensci/drake"
)
if (FALSE) { # \dontrun{
issue_authors_tbl <- update_issue_authors (repo_urls, "path/to/repo-data-out")
} # }