You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas遍历repo列遇KeyError,求解决Selenium打开GitHub链接问题

问题解决

错误原因

你代码里的row['repo_name']是硬编码的字符串,而repo_name是循环变量,实际代表repo1、repo2这类列名。DataFrame中不存在名为repo_name的列,因此触发KeyError。另外原代码未处理NaN值,直接访问会生成包含无效内容的URL。

修正后的代码

import pandas as pd
import time
# 假设driver已提前初始化完成

repo_cols = [col for col in df1.columns if 'repo' in col]

for index, row in df1.iterrows():
    user = row['user']
    for repo_col in repo_cols:
        repo = row[repo_col]
        # 跳过空仓库名,避免无效URL访问
        if pd.notna(repo):
            current_url = f'https://github.com/{user}/{repo}/graphs/contributors'
            driver.get(current_url)
            time.sleep(0.5)

关键修改点

  • 将row['repo_name']改为row[repo_col](变量无需用引号包裹,直接引用即可获取对应列的内容)
  • 增加pd.notna(repo)判断,自动跳过NaN的仓库条目,避免访问无效链接

内容的提问来源于stack exchange,提问作者Fred

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 01:10:55