GitHub主题搜索仓库爬虫问题:全量爬取失败等三类异常
解决GitHub仓库爬虫的三类问题
问题1:仅能爬取少量仓库,无法获取全量结果
GitHub网页搜索有硬性限制:最多显示100页(每页10个仓库,共1000个),且频繁请求会触发反爬机制导致IP被封。如果要获取超过1000个仓库数据,必须改用GitHub Search API——API支持最多1000条结果(per_page设为100,共10页),若需更多,可通过按创建时间/更新时间分段搜索实现。同时使用个人访问令牌(token)可提升API请求限额。
问题2:about字段为空导致DataFrame构建报错
原代码单独提取所有about字段,忽略了“部分仓库无about”的情况,导致stock_about长度与stock_names、stock_urls不一致,触发ValueError: All arrays must be of the same length。
解决方案:先定位每个仓库的父容器,对每个仓库单独提取信息——有about则取文本,无则填充空字符串,确保三个列表长度完全匹配。
问题3:仓库所有者名称提取错误
原正则表达式re.sub(r"\/(.*)\/(.*)", "\1", link)存在转义错误(应使用r"\1"),且匹配逻辑不严谨。更简单可靠的方式是通过字符串分割:仓库链接格式为/owner/repo,直接split('/')取第二个元素即可。
修正后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } # 若需获取更多数据,可改用GitHub API(需替换为自己的token) # GITHUB_TOKEN = "your_personal_access_token" # api_url = "https://api.github.com/search/repositories?q=nlp&per_page=100" # headers['Authorization'] = f'token {GITHUB_TOKEN}' search_base = "https://github.com/search?p=" query = "nlp&type=Repositories" repos_data = [] # 网页搜索最多爬100页(GitHub官方限制) for page in range(1, 101): req = requests.get(f"{search_base}{page}&q={query}", headers=headers) soup = BeautifulSoup(req.text, "html.parser") # 定位每个仓库的容器节点 repo_containers = soup.select("div.Box-row") for container in repo_containers: # 提取仓库链接和所有者名称 repo_link_elem = container.select_one("a.v-align-middle") if repo_link_elem: repo_href = repo_link_elem['href'] repo_url = f"https://github.com{repo_href}" owner_name = repo_href.split('/')[1] # 直接分割取所有者名称 else: repo_url = "" owner_name = "" # 提取about信息,无则填空字符串 about_elem = container.select_one("p.mb-1") about_text = about_elem.text.strip() if about_elem else "" # 将单条仓库数据存入列表 repos_data.append({ "name": owner_name, "url": repo_url, "about": about_text }) # 直接从结构化列表构建DataFrame,彻底避免长度不匹配问题 df = pd.DataFrame(repos_data) print(df.head())
关键改进说明
- 按仓库容器遍历:确保每条仓库数据的三个字段一一对应,即使about为空也填充空字符串,彻底解决长度不匹配问题。
- 所有者名称提取优化:用
split('/')[1]替代正则,逻辑更清晰、不易出错。 - 分页限制处理:网页搜索最多爬100页,若需更多数据,注释部分提供了API调用示例,需自行申请GitHub个人访问令牌。
内容的提问来源于stack exchange,提问作者ladybug
相关产品推荐
相关产品推荐

