GitHub用户月度贡献爬取无输出问题排查求助
GitHub用户月度贡献数据爬取代码无输出问题排查及修复
核心问题及解决办法
无差别异常捕获掩盖所有错误
代码中except:直接跳过所有异常,包括用户获取失败、Selenium语法错误、元素未找到、页面加载失败等,导致你完全无法定位问题根源。
修复:捕获特定异常并打印错误信息,示例:except Exception as e: print(f"处理用户{user} {y}_{m}时出错: {e}") continueSelenium语法错误(Java写法误用)
Python版Selenium方法为小写开头,代码中driver.findElement()、getText()是Java语法,Python中应改为driver.find_element()和.text属性:# 错误写法 # contribution = driver.findElement(webdriver.By.xpath("...")).getText(); # 正确写法 from selenium.webdriver.common.by import By contribution = driver.find_element(By.XPATH, "//*[@id='js-contribution-activity']/div/div/div/div/details/summary/span[1]").text等待时间过短,页面未渲染完成
0.1秒的sleep远不足以加载GitHub贡献页面,元素尚未渲染就去查找会触发异常。建议用显式等待替代固定sleep:from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 等待元素出现,最多等10秒 element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//*[@id='js-contribution-activity']/div/div/div/div/details/summary/span[1]")) ) contribution = element.text用户获取失败后仍执行无效循环
当row['user']不存在时,user变量未赋值,后续请求的URL会是无效的https://github.com/?tab=...,但异常被跳过,导致无数据写入。
修复:获取user失败直接跳过当前行:try: user = row['user'] except KeyError: print(f"第{index}行无user字段,跳过") continue循环范围未覆盖"至今"
代码中年份循环到2022(range(2004,2023)为左闭右开),若当前年份大于2022会漏掉后续数据。动态获取当前年份和月份:from datetime import datetime now = datetime.now() current_year = now.year current_month = now.month for y in range(2004, current_year + 1): # 当年份是当前年时,只循环到当前月份 end_month = current_month if y == current_year else 12 for m in range(1, end_month + 1): # 后续逻辑 pass
修复后示例代码片段
import time import pandas as pd from datetime import datetime from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, NoSuchElementException df1 = pd.read_csv("你的原始CSV路径.csv") driver = webdriver.Chrome() # 确保ChromeDriver配置正确 now = datetime.now() current_year = now.year current_month = now.month for index, row in df1.iterrows(): try: user = row['user'] except KeyError: print(f"第{index}行无user字段,跳过") continue for y in range(2004, current_year + 1): end_month = current_month if y == current_year else 12 for m in range(1, end_month + 1): try: current_url = f'https://github.com/{user}?tab=overview&from={y}-{m:02d}-01&to={y}-{m:02d}-31' print(f"正在爬取: {current_url}") driver.get(current_url) # 显式等待元素加载 element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//*[@id='js-contribution-activity']/div/div/div/div/details/summary/span[1]")) ) contribution = element.text # 格式化月份为两位数字,比如2022_01而非2022_1 df1.loc[index, f'{y}_{m:02d}'] = contribution except (TimeoutException, NoSuchElementException) as e: print(f"爬取{user} {y}_{m}失败: {e}") df1.loc[index, f'{y}_{m:02d}'] = "无数据" except Exception as e: print(f"未知错误: {e}") continue driver.quit() print(df1) df1.to_csv('C:/Users/fredr/Desktop/output today.csv', index=False)
内容的提问来源于stack exchange,提问作者KDWB
相关产品推荐
相关产品推荐

