You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python/Shell脚本批量统计指定用户的GitHub PR搜索结果数

批量统计GitHub用户符合指定条件的PR数量(无API版)

已知手动在GitHub搜索框执行is:pr archived:false is:closed base:master base:main merged:>=2023-01-01 author:<用户名>可获取单个用户的目标PR结果,以下提供两种无API依赖的批量统计方案:

Python 实现

通过请求GitHub搜索页面并解析HTML提取结果计数,需依赖requests和beautifulsoup4库。

代码示例

import requests
from bs4 import BeautifulSoup
import sys
import time

def count_prs(author):
    search_query = f"is:pr archived:false is:closed base:master base:main merged:>=2023-01-01 author:{author}"
    encoded_query = requests.utils.quote(search_query)
    url = f"https://github.com/search?q={encoded_query}&type=pullrequests"
    
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    try:
        response = requests.get(url, headers=headers)
        response.raise_for_status()
    except requests.exceptions.RequestException:
        return -1  # 标记请求失败
    
    soup = BeautifulSoup(response.text, "html.parser")
    result_elem = soup.find("div", class_="search-title")
    if not result_elem:
        return 0
    
    count_text = result_elem.get_text(strip=True).split()[0].replace(",", "")
    return int(count_text) if count_text.isdigit() else 0

if __name__ == "__main__":
    if len(sys.argv) < 2:
        print("用法: python count_prs.py <用户名1> <用户名2> ...")
        sys.exit(1)
    
    for author in sys.argv[1:]:
        count = count_prs(author)
        if count == -1:
            print(f"{author}: 请求失败")
        else:
            print(f"{author}: {count} 符合条件的PR")
        time.sleep(2)  # 避免触发反爬限制

使用说明

  1. 安装依赖:pip install requests beautifulsoup4
  2. 运行脚本:python count_prs.py author1 author2 ...
  3. 注意事项:
    • 搜索结果超过1000时,GitHub会显示"1,000+ results",脚本将返回1000
    • 不要移除延迟,否则可能被GitHub临时封禁IP

Shell 实现

用curl获取页面,通过正则匹配提取结果数,无需额外安装库(依赖系统自带工具)。

代码示例

#!/bin/bash

if [ $# -eq 0 ]; then
    echo "用法: ./count_prs.sh <用户名1> <用户名2> ..."
    exit 1
fi

UA="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"

for author in "$@"; do
    # URL编码搜索查询
    raw_query="is:pr archived:false is:closed base:master base:main merged:>=2023-01-01 author:$author"
    encoded_query=$(printf "%s" "$raw_query" | xxd -plain | tr -d '\n' | sed 's/\(..\)/%\1/g')
    search_url="https://github.com/search?q=$encoded_query&type=pullrequests"
    
    # 获取结果计数
    count=$(curl -s -A "$UA" "$search_url" | grep -oP '<div class="search-title">.*?(\d{1,3}(?:,\d{3})*) results?' | sed -E 's/.*?(\d{1,3}(?:,\d{3})*).*/\1/' | tr -d ',')
    
    # 处理无结果或请求失败的情况
    if [ -z "$count" ]; then
        count=0
    fi
    
    echo "$author: $count 符合条件的PR"
    sleep 2  # 延迟防反爬
done

使用说明

  1. 保存脚本为count_prs.sh,赋予执行权限:chmod +x count_prs.sh
  2. 运行脚本:./count_prs.sh author1 author2 ...
  3. 注意事项:
    • 确保系统安装了curl、grep(支持Perl正则)、sed、xxd
    • 结果超过1000时仅返回1000,无法获取准确总数

内容的提问来源于stack exchange,提问作者Joseph Salom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 20:15:39