使用Python的Beautiful Soup爬取GitHub commits时返回None如何解决
问题原因
- 直接发起无请求头的HTTP请求,触发GitHub反爬策略,返回的HTML内容不完整,缺少目标节点
- 所用CSS选择器匹配的是GitHub旧版页面结构,当前平台页面结构已更新,旧选择器无法定位到commit计数元素
修正代码
from bs4 import BeautifulSoup import requests # 添加浏览器请求头绕过基础反爬 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } html = requests.get('https://github.com/pnp/cli-microsoft365', headers=headers).text soup = BeautifulSoup(html, 'html.parser') # 适配新版页面的选择器,新增判空逻辑避免异常 commit_count_node = soup.select_one('a[href*="/commits"] strong') if commit_count_node: commit_count = commit_count_node.text.strip() print(commit_count) else: print("未获取到commit数据")
优化建议
如果需要长期稳定获取这类数据,建议使用GitHub官方提供的API接口,无需解析前端页面,不会受页面结构迭代影响,返回的结构化数据也更便于后续处理。
内容的提问来源于stack exchange,提问作者lex
相关产品推荐
相关产品推荐

