GitHub仓库更新检测爬虫随机失效问题求助
GitHub仓库更新检测爬虫间歇性失效问题
我正在开发一个轻量网页爬虫,用来检测GitHub仓库是否更新并发送通知。目前遇到的问题是:爬虫有时能正常工作,通过BeautifulSoup获取正确HTML并找到relative-time/datetime标签(我的仓库配置里这类标签至少有3个);但有时运行完全相同的代码却拿不到预期的HTML,怀疑是认证或会话超时导致的。
我的需求是尽量避免手动提取Cookie/Header嵌入脚本,希望用户只需要输入账号密码就能用,除非万不得已不想手动处理所有Cookie和Header。
import requests from bs4 import BeautifulSoup import datetime #Return a Soup Value using a Response def initialize_soup(resp): soup = BeautifulSoup(resp.text, 'html.parser') return soup #Get the most recent date the repository was updated def get_last_updated(bs): #look for relative-time tag time_var = bs.find('relative-time') string = str(time_var) #isolate down to datetime param s = string.find("datetime") value = s #Grab static values to create a date-time value date_string = string[value + 10:(value + 10 + 10)] time_string = string[value + 21:(value + 21 + 5)] #Concatentate into new value last_updated_string = date_string + ' ' + time_string + ':00' #If the web scraper bugs out, it'll return a 'default' datetime so that code doesn't #break if len(last_updated_string) < 15: last_updated_string = "1900-01-01 12:00:00" print("CONNECTION ERROR") last_updated = datetime.datetime.strptime(last_updated_string, "%Y-%m-%d %H:%M:%S") return last_updated #compare dates to see if GitHub has updated def is_updated(old_datetime, new_date_time): if old_datetime < new_date_time: print("GitHub updated") else: print("Not Updated") login_url = "https://github.com/session" login = 'login value omitted here' password = 'password omitted here' with requests.session() as s: req = s.get(login_url).text html = BeautifulSoup(req,"html.parser") token = html.find("input", {"name": "authenticity_token"}). attrs["value"] time = html.find("input", {"name": "timestamp"}).attrs["value"] timeSecret = html.find("input", {"name": "timestamp_secret"}). attrs["value"] payload = { "authenticity_token": token, "login": login, "password": password, "timestamp": time, "timestamp_secret": timeSecret } res = s.post(login_url, data = payload) repository_url = "working repository link" #get the first date time value (we'll use this to compare dates on an interval) r = s.get(repository_url) bs = initialize_soup(r) prev_dt = get_last_updated(bs) ## These lines just test to see if the code is working/when it's not working #print(bs.find_all("datetime")) #print(bs) #On an interval, check the repository's last updated date and compare against old date. import time while True: new_response = s.get(repository_url) new_soup = initialize_soup(new_response) new_dt = get_last_updated(new_soup) if prev_dt < new_dt: print("There's an update on GitHub waiting for you!") #if we have a new update to report, update the values prev_dt = new_dt else: print("Not updated -- Testing Purposes Only") time.sleep(30)
更新1:该问题似乎和GitHub的请求调用次数有关,推测会话(Sessions)存在异常。我已经引入Selenium优化,但还是偶尔出现连接问题,求可靠解决方案。
更新2:我查看了报错时返回的HTML,发现GitHub页面返回错误信息:Failed to load latest commit information。
问题成因
- 反爬机制触发:频繁的无浏览器特征请求会被GitHub判定为自动化工具,限制访问,导致提交信息加载失败。
- 会话失效:requests的会话可能因长时间闲置或GitHub的会话过期策略导致认证失效,无法获取需权限的页面内容。
- 动态内容渲染:GitHub部分页面内容通过JavaScript动态生成,纯requests+BeautifulSoup无法捕获动态加载的
relative-time标签。
解决方案
1. 改用GitHub API(最优方案)
完全避开网页爬取的问题,API接口稳定且认证简单:
- 在GitHub设置中生成个人访问令牌(PAT),权限仅需勾选
repo即可。 - 调用
GET /repos/{owner}/{repo}接口,返回的updated_at字段直接是仓库最后更新时间,无需解析HTML。
示例代码:
import requests import time from datetime import datetime GITHUB_TOKEN = "你的个人访问令牌" REPO_API_URL = "https://api.github.com/repos/{owner}/{repo}" # 替换为实际仓库路径 headers = {"Authorization": f"token {GITHUB_TOKEN}"} def get_last_updated(): resp = requests.get(REPO_API_URL, headers=headers) resp.raise_for_status() # 捕获请求错误 data = resp.json() # 转换ISO格式时间为datetime对象 return datetime.fromisoformat(data["updated_at"].replace("Z", "+00:00")) prev_dt = get_last_updated() while True: new_dt = get_last_updated() if prev_dt < new_dt: print("GitHub仓库有更新!") prev_dt = new_dt else: print("暂无更新 -- 测试用") time.sleep(300) # 建议拉长间隔,避免API限流
2. 优化Selenium方案(坚持爬网页时用)
- 替换为
undetected-chromedriver,绕过GitHub的自动化检测。 - 添加随机请求间隔、页面滚动等真实浏览器行为。
- 持久化浏览器Cookie,避免重复登录。
示例调整:
from selenium import webdriver from selenium.webdriver.common.by import By import random import time # 初始化undetected-chromedriver(需先安装:pip install undetected-chromedriver) driver = webdriver.Chrome() driver.get("https://github.com/login") # 自动填充登录信息(若有两步验证需手动处理) driver.find_element(By.ID, "login_field").send_keys("你的账号") driver.find_element(By.ID, "password").send_keys("你的密码") driver.find_element(By.NAME, "commit").click() repo_url = "你的仓库链接" driver.get(repo_url) # 获取初始更新时间 prev_dt = driver.find_element(By.TAG_NAME, "relative-time").get_attribute("datetime") while True: # 随机间隔60-120秒 time.sleep(random.randint(60, 120)) driver.refresh() try: new_dt = driver.find_element(By.TAG_NAME, "relative-time").get_attribute("datetime") if prev_dt < new_dt: print("GitHub仓库有更新!") prev_dt = new_dt else: print("暂无更新 -- 测试用") except Exception as e: print("加载失败,重新刷新页面") driver.refresh()
3. 修复原有requests方案(临时应急)
- 添加浏览器请求头,模拟真实用户访问:
headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } - 每次请求前检查会话有效性,若失效则重新登录。
- 拉长请求间隔至60秒以上,避免触发频率限制。
注意事项
- GitHub API对认证用户的速率限制为每小时5000次请求,完全满足检测需求。
- 爬网页时务必遵守GitHub服务条款,避免过度请求导致账号受限。
内容的提问来源于stack exchange,提问作者TripleCute
相关产品推荐
相关产品推荐

