如何突破GitHub Commits API限制获取全部相关提交记录?
GitHub搜索API获取全部提交记录的解决方案
GitHub的Commit搜索API存在最多返回1000条结果的硬限制,这是官方设定的规则,无法通过调整分页参数突破。要获取全部14万+条结果,可通过以下两种方式实现:
1. 按时间范围拆分查询
利用GitHub搜索的时间过滤参数created,将整个时间区间拆分为多个小时间段,确保每个子查询的结果数不超过1000条,再分别查询后合并结果。
示例R代码调整:
library(jsonlite) library(httpuv) library(httr) # 初始化OAuth(保留你原有的认证代码) oauth_endpoints("github") myapp <- oauth_app(appname = "name", key = "insert_key", secret = "insert_secret") github_token <- oauth2.0_token(oauth_endpoints("github"), myapp) gtoken <- config(token = github_token) # 初始化结果数据集 df_all <- data.frame(matrix(nrow = 0, ncol = 4)) names(df_all) = c("items_url", "items_sha", "items_commit_url", "items_commit_message") # 拆分时间区间为年度(可根据实际结果数调整为月度/周度) years <- 2010:2024 for (year in years) { start_date <- paste0(year, "-01-01") end_date <- paste0(year + 1, "-01-01") # 每个时间区间内最多查询10页(1000条) for (page in 1:10) { # 构造带时间过滤的查询URL getURL <- paste0( "https://api.github.com/search/commits?q=%22sentiment%22+created:>==", start_date, "+created:<", end_date, "&per_page=100&page=", page ) req <- GET(getURL, gtoken) # 检查请求是否成功 if (status_code(req) != 200) { warning(paste("请求失败:年份", year, "页码", page)) next } char <- rawToChar(req$content) df <- jsonlite::fromJSON(char) # 当前页无结果时,终止该时间区间的查询 if (length(df$items) == 0) break df_iteration <- data.frame( items_url = df$items$url, items_sha = df$items$sha, items_commit_url = df$items$commit$url, items_commit_message = df$items$commit$message ) df_all <- rbind(df_all, df_iteration) # 添加请求间隔,避免触发速率限制 Sys.sleep(1) } } # 去重(基于sha字段) df_all <- df_all[!duplicated(df_all$items_sha), ]
如果某个时间区间的结果仍超过1000条,可进一步缩小时间粒度(比如拆分为月度)。
2. 使用GitHub GraphQL API
GraphQL API支持基于cursor的滚动分页,可遍历更多结果(仍需遵守速率限制)。你可以用R的ghql包实现:
核心思路:
- 构造GraphQL查询,每次获取100条结果,并返回下一页的
cursor - 循环调用直到
hasNextPage为false
示例代码片段:
library(ghql) library(jsonlite) # 初始化客户端 con <- GraphqlClient$new( url = "https://api.github.com/graphql", headers = list(Authorization = paste0("token ", github_token$credentials$access_token)) ) # 定义查询模板 query_template <- ' query ($cursor: String) { search(query: "\"sentiment\" type:commit", type: COMMIT, first: 100, after: $cursor) { pageInfo { endCursor hasNextPage } edges { node { url oid commit { url message } } } } } ' q <- Query$new()$query("getCommits", query_template) df_all <- data.frame() has_next <- TRUE cursor <- NULL while (has_next) { result <- con$exec(q, variables = list(cursor = cursor)) result_json <- fromJSON(result) # 提取当前页数据 current_page <- data.frame( items_url = result_json$data$search$edges$node$url, items_sha = result_json$data$search$edges$node$oid, items_commit_url = result_json$data$search$edges$node$commit$url, items_commit_message = result_json$data$search$edges$node$commit$message ) df_all <- rbind(df_all, current_page) # 更新cursor和循环状态 has_next <- result_json$data$search$pageInfo$hasNextPage cursor <- result_json$data$search$pageInfo$endCursor Sys.sleep(1) }
关键注意事项
- 速率限制:认证用户每小时最多5000次请求,需添加
Sys.sleep()控制请求间隔,避免被限制。 - 结果去重:不同查询可能返回重复提交,务必通过
sha(GraphQL中为oid)字段去重。
内容的提问来源于stack exchange,提问作者Erik Brole
相关产品推荐
相关产品推荐

