You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GitHub用户月度贡献爬取无输出问题排查求助

GitHub用户月度贡献数据爬取代码无输出问题排查及修复

核心问题及解决办法

  • 无差别异常捕获掩盖所有错误
    代码中except:直接跳过所有异常,包括用户获取失败、Selenium语法错误、元素未找到、页面加载失败等,导致你完全无法定位问题根源。
    修复:捕获特定异常并打印错误信息,示例:

    except Exception as e:
        print(f"处理用户{user} {y}_{m}时出错: {e}")
        continue
    
  • Selenium语法错误(Java写法误用)
    Python版Selenium方法为小写开头,代码中driver.findElement()、getText()是Java语法,Python中应改为driver.find_element()和.text属性:

    # 错误写法
    # contribution = driver.findElement(webdriver.By.xpath("...")).getText();
    # 正确写法
    from selenium.webdriver.common.by import By
    contribution = driver.find_element(By.XPATH, "//*[@id='js-contribution-activity']/div/div/div/div/details/summary/span[1]").text
    
  • 等待时间过短,页面未渲染完成
    0.1秒的sleep远不足以加载GitHub贡献页面,元素尚未渲染就去查找会触发异常。建议用显式等待替代固定sleep:

    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    # 等待元素出现,最多等10秒
    element = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//*[@id='js-contribution-activity']/div/div/div/div/details/summary/span[1]"))
    )
    contribution = element.text
    
  • 用户获取失败后仍执行无效循环
    当row['user']不存在时,user变量未赋值,后续请求的URL会是无效的https://github.com/?tab=...,但异常被跳过,导致无数据写入。
    修复:获取user失败直接跳过当前行:

    try:
        user = row['user']
    except KeyError:
        print(f"第{index}行无user字段,跳过")
        continue
    
  • 循环范围未覆盖"至今"
    代码中年份循环到2022(range(2004,2023)为左闭右开),若当前年份大于2022会漏掉后续数据。动态获取当前年份和月份:

    from datetime import datetime
    now = datetime.now()
    current_year = now.year
    current_month = now.month
    
    for y in range(2004, current_year + 1):
        # 当年份是当前年时,只循环到当前月份
        end_month = current_month if y == current_year else 12
        for m in range(1, end_month + 1):
            # 后续逻辑
            pass
    

修复后示例代码片段

import time
import pandas as pd
from datetime import datetime
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, NoSuchElementException

df1 = pd.read_csv("你的原始CSV路径.csv")
driver = webdriver.Chrome()  # 确保ChromeDriver配置正确
now = datetime.now()
current_year = now.year
current_month = now.month

for index, row in df1.iterrows():
    try:
        user = row['user']
    except KeyError:
        print(f"第{index}行无user字段,跳过")
        continue
    
    for y in range(2004, current_year + 1):
        end_month = current_month if y == current_year else 12
        for m in range(1, end_month + 1):
            try:
                current_url = f'https://github.com/{user}?tab=overview&from={y}-{m:02d}-01&to={y}-{m:02d}-31'
                print(f"正在爬取: {current_url}")
                driver.get(current_url)
                
                # 显式等待元素加载
                element = WebDriverWait(driver, 10).until(
                    EC.presence_of_element_located((By.XPATH, "//*[@id='js-contribution-activity']/div/div/div/div/details/summary/span[1]"))
                )
                contribution = element.text
                # 格式化月份为两位数字,比如2022_01而非2022_1
                df1.loc[index, f'{y}_{m:02d}'] = contribution
                
            except (TimeoutException, NoSuchElementException) as e:
                print(f"爬取{user} {y}_{m}失败: {e}")
                df1.loc[index, f'{y}_{m:02d}'] = "无数据"
            except Exception as e:
                print(f"未知错误: {e}")
                continue

driver.quit()
print(df1)
df1.to_csv('C:/Users/fredr/Desktop/output today.csv', index=False)

内容的提问来源于stack exchange,提问作者KDWB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 16:01:12