You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GitHub仓库更新检测爬虫随机失效问题求助

GitHub仓库更新检测爬虫间歇性失效问题

我正在开发一个轻量网页爬虫,用来检测GitHub仓库是否更新并发送通知。目前遇到的问题是:爬虫有时能正常工作,通过BeautifulSoup获取正确HTML并找到relative-time/datetime标签(我的仓库配置里这类标签至少有3个);但有时运行完全相同的代码却拿不到预期的HTML,怀疑是认证或会话超时导致的。

我的需求是尽量避免手动提取Cookie/Header嵌入脚本,希望用户只需要输入账号密码就能用,除非万不得已不想手动处理所有Cookie和Header。

import requests
from bs4 import BeautifulSoup
import datetime


#Return a Soup Value using a Response
def initialize_soup(resp):
    soup = BeautifulSoup(resp.text, 'html.parser')
    return soup

#Get the most recent date the repository was updated
def get_last_updated(bs):

    #look for relative-time tag
    time_var = bs.find('relative-time')
    string = str(time_var)

    #isolate down to datetime param
    s = string.find("datetime")
    value = s

    #Grab static values to create a date-time value
    date_string = string[value + 10:(value + 10 + 10)]
    time_string = string[value + 21:(value + 21 + 5)]
    
    #Concatentate into new value
    last_updated_string = date_string + ' ' + time_string + ':00'

    #If the web scraper bugs out, it'll return a 'default' datetime so that code doesn't 
    #break
    if len(last_updated_string) < 15:
        last_updated_string = "1900-01-01 12:00:00"
        print("CONNECTION ERROR")

    last_updated = datetime.datetime.strptime(last_updated_string, "%Y-%m-%d %H:%M:%S")
    return last_updated

#compare dates to see if GitHub has updated
def is_updated(old_datetime, new_date_time):
    if old_datetime < new_date_time:
        print("GitHub updated")
    else:
        print("Not Updated")



login_url = "https://github.com/session"

login = 'login value omitted here'
password = 'password omitted here'

with requests.session() as s:
    req = s.get(login_url).text
    html = BeautifulSoup(req,"html.parser")
    token = html.find("input", {"name": "authenticity_token"}). attrs["value"]
    time = html.find("input", {"name": "timestamp"}).attrs["value"]
    timeSecret = html.find("input", {"name": "timestamp_secret"}). attrs["value"]
payload = {
    "authenticity_token": token,
    "login": login,
    "password": password,
    "timestamp": time,
    "timestamp_secret": timeSecret
}

res = s.post(login_url, data = payload)
repository_url = "working repository link"

#get the first date time value (we'll use this to compare dates on an interval)
r = s.get(repository_url)
bs = initialize_soup(r)
prev_dt = get_last_updated(bs)

## These lines just test to see if the code is working/when it's not working
#print(bs.find_all("datetime"))
#print(bs)

#On an interval, check the repository's last updated date and compare against old date.
import time
while True:
    new_response = s.get(repository_url)
    new_soup = initialize_soup(new_response)
    new_dt = get_last_updated(new_soup)

    if prev_dt < new_dt:
        print("There's an update on GitHub waiting for you!")
        #if we have a new update to report, update the values
        prev_dt = new_dt
    else:
        print("Not updated -- Testing Purposes Only")
    time.sleep(30)

更新1:该问题似乎和GitHub的请求调用次数有关,推测会话(Sessions)存在异常。我已经引入Selenium优化,但还是偶尔出现连接问题,求可靠解决方案。

更新2:我查看了报错时返回的HTML,发现GitHub页面返回错误信息:Failed to load latest commit information。


问题成因

  1. 反爬机制触发:频繁的无浏览器特征请求会被GitHub判定为自动化工具,限制访问,导致提交信息加载失败。
  2. 会话失效:requests的会话可能因长时间闲置或GitHub的会话过期策略导致认证失效,无法获取需权限的页面内容。
  3. 动态内容渲染:GitHub部分页面内容通过JavaScript动态生成,纯requests+BeautifulSoup无法捕获动态加载的relative-time标签。

解决方案

1. 改用GitHub API(最优方案)

完全避开网页爬取的问题,API接口稳定且认证简单:

  • 在GitHub设置中生成个人访问令牌(PAT),权限仅需勾选repo即可。
  • 调用GET /repos/{owner}/{repo}接口,返回的updated_at字段直接是仓库最后更新时间,无需解析HTML。

示例代码:

import requests
import time
from datetime import datetime

GITHUB_TOKEN = "你的个人访问令牌"
REPO_API_URL = "https://api.github.com/repos/{owner}/{repo}"  # 替换为实际仓库路径

headers = {"Authorization": f"token {GITHUB_TOKEN}"}

def get_last_updated():
    resp = requests.get(REPO_API_URL, headers=headers)
    resp.raise_for_status()  # 捕获请求错误
    data = resp.json()
    # 转换ISO格式时间为datetime对象
    return datetime.fromisoformat(data["updated_at"].replace("Z", "+00:00"))

prev_dt = get_last_updated()
while True:
    new_dt = get_last_updated()
    if prev_dt < new_dt:
        print("GitHub仓库有更新!")
        prev_dt = new_dt
    else:
        print("暂无更新 -- 测试用")
    time.sleep(300)  # 建议拉长间隔,避免API限流

2. 优化Selenium方案(坚持爬网页时用)

  • 替换为undetected-chromedriver,绕过GitHub的自动化检测。
  • 添加随机请求间隔、页面滚动等真实浏览器行为。
  • 持久化浏览器Cookie,避免重复登录。

示例调整:

from selenium import webdriver
from selenium.webdriver.common.by import By
import random
import time

# 初始化undetected-chromedriver(需先安装:pip install undetected-chromedriver)
driver = webdriver.Chrome()
driver.get("https://github.com/login")

# 自动填充登录信息(若有两步验证需手动处理)
driver.find_element(By.ID, "login_field").send_keys("你的账号")
driver.find_element(By.ID, "password").send_keys("你的密码")
driver.find_element(By.NAME, "commit").click()

repo_url = "你的仓库链接"
driver.get(repo_url)
# 获取初始更新时间
prev_dt = driver.find_element(By.TAG_NAME, "relative-time").get_attribute("datetime")

while True:
    # 随机间隔60-120秒
    time.sleep(random.randint(60, 120))
    driver.refresh()
    try:
        new_dt = driver.find_element(By.TAG_NAME, "relative-time").get_attribute("datetime")
        if prev_dt < new_dt:
            print("GitHub仓库有更新!")
            prev_dt = new_dt
        else:
            print("暂无更新 -- 测试用")
    except Exception as e:
        print("加载失败,重新刷新页面")
        driver.refresh()

3. 修复原有requests方案(临时应急)

  • 添加浏览器请求头,模拟真实用户访问:
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    
  • 每次请求前检查会话有效性,若失效则重新登录。
  • 拉长请求间隔至60秒以上,避免触发频率限制。

注意事项

  • GitHub API对认证用户的速率限制为每小时5000次请求,完全满足检测需求。
  • 爬网页时务必遵守GitHub服务条款,避免过度请求导致账号受限。

内容的提问来源于stack exchange,提问作者TripleCute

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 06:42:07