Github仓库爬虫返回空DataFrame问题排查求助
Github搜索页爬虫返回空DataFrame的问题修复
问题根源
你代码里用的类名(比如Box-sc-g0xbh4-0、Link__StyledLink-sc-14289xe-0)是Github动态生成的随机类名,这类类名会随页面更新、加载会话变化,导致BeautifulSoup无法匹配到目标元素,最终所有列表为空,生成空DataFrame。
修正方案
改用语义化、稳定的选择器定位元素,这类选择器不会随页面动态更新变化:
import requests from bs4 import BeautifulSoup import pandas as pd rep_link = [] description = [] date = [] url = "https://github.com/search?q=excel+import&type=repositories" r = requests.get(url) soup = BeautifulSoup(r.text, features="html.parser") # 遍历每个仓库条目(用稳定的Box-row类) for item in soup.find_all('div', class_='Box-row'): # 获取仓库名称 repo_a = item.find('a', class_='v-align-middle') rep_link.append(repo_a.get_text(strip=True) if repo_a else None) # 获取项目描述 desc_p = item.find('p', class_='col-9') description.append(desc_p.get_text(strip=True) if desc_p else None) # 获取最后更新日期 time_tag = item.find('relative-time') date.append(time_tag.get_text(strip=True) if time_tag else None) # 生成DataFrame并过滤无效条目 df = pd.DataFrame({ "Repo": rep_link, "Description": description, "Date": date }).dropna(subset=["Repo"]) print(df)
额外说明
- 加入空值判断,避免部分仓库无描述等情况导致代码报错
- 用
get_text(strip=True)去除文本前后空白字符,优化数据整洁度 - 最后过滤掉无仓库名的无效条目,保证DataFrame有效性
内容的提问来源于stack exchange,提问作者Baha Dawood ud-Din Rehman
相关产品推荐
相关产品推荐

