You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Github仓库爬虫返回空DataFrame问题排查求助

Github搜索页爬虫返回空DataFrame的问题修复

问题根源

你代码里用的类名(比如Box-sc-g0xbh4-0、Link__StyledLink-sc-14289xe-0)是Github动态生成的随机类名,这类类名会随页面更新、加载会话变化,导致BeautifulSoup无法匹配到目标元素,最终所有列表为空,生成空DataFrame。

修正方案

改用语义化、稳定的选择器定位元素,这类选择器不会随页面动态更新变化:

import requests
from bs4 import BeautifulSoup
import pandas as pd

rep_link = []
description = []
date = []
url = "https://github.com/search?q=excel+import&type=repositories"

r = requests.get(url)
soup = BeautifulSoup(r.text, features="html.parser")

# 遍历每个仓库条目(用稳定的Box-row类)
for item in soup.find_all('div', class_='Box-row'):
    # 获取仓库名称
    repo_a = item.find('a', class_='v-align-middle')
    rep_link.append(repo_a.get_text(strip=True) if repo_a else None)
    
    # 获取项目描述
    desc_p = item.find('p', class_='col-9')
    description.append(desc_p.get_text(strip=True) if desc_p else None)
    
    # 获取最后更新日期
    time_tag = item.find('relative-time')
    date.append(time_tag.get_text(strip=True) if time_tag else None)

# 生成DataFrame并过滤无效条目
df = pd.DataFrame({
    "Repo": rep_link,
    "Description": description,
    "Date": date
}).dropna(subset=["Repo"])

print(df)

额外说明

  • 加入空值判断,避免部分仓库无描述等情况导致代码报错
  • 用get_text(strip=True)去除文本前后空白字符,优化数据整洁度
  • 最后过滤掉无仓库名的无效条目,保证DataFrame有效性

内容的提问来源于stack exchange,提问作者Baha Dawood ud-Din Rehman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 19:38:21