如何使用Python和BeautifulSoup精准获取GitHub仓库的提交次数?
解决GitHub仓库提交数爬取不稳定的问题
这个问题我之前也碰到过!核心原因就是你用的d-none d-sm-inline是GitHub页面里的通用响应式类,很多元素都会用到,所以抓的时候很容易误匹配。给你几个靠谱的解决方案,帮你精准定位到提交数:
方法1:利用Git Stats模块的上下文定位
提交数是在标注为"Git stats"的模块下,我们可以先找到这个模块的标题,再顺着结构找到对应的提交数字:
import requests from bs4 import BeautifulSoup as bs r = requests.get(source_code_link) soup = bs(r.content, 'lxml') # 先定位到Git stats的隐藏标题 git_stats_title = soup.find('h2', class_='sr-only', string='Git stats') if git_stats_title: # 找到标题后面的ul列表,再取第一个li里的strong commit_strong = git_stats_title.find_next('ul').find('li').find('strong') if commit_strong: commit_count = commit_strong.text.strip() print(f"提交次数:{commit_count}")
这种方式依赖页面结构的关联性,能确保我们只在Git stats模块里找数据,不会被其他区域的相同class干扰。
方法2:利用aria-label属性精准匹配
提交数旁边的span带有aria-label="Commits on master"这样的属性,我们可以通过这个属性来定位:
import requests from bs4 import BeautifulSoup as bs r = requests.get(source_code_link) soup = bs(r.content, 'lxml') # 找到aria-label包含"Commits"的span元素 commit_label_span = soup.find('span', attrs={'aria-label': lambda x: x and 'Commits' in x}) if commit_label_span: # 这个span和strong标签在同一个父span里,直接找父元素下的strong commit_strong = commit_label_span.parent.find('strong') if commit_strong: commit_count = commit_strong.text.strip() print(f"提交次数:{commit_count}")
这种方法不依赖层级结构,只要标签的aria-label包含提交相关描述就能匹配,兼容性更强。
方法3:使用精确的CSS选择器
直接用CSS选择器组合层级和属性,一步定位到目标元素:
import requests from bs4 import BeautifulSoup as bs r = requests.get(source_code_link) soup = bs(r.content, 'lxml') # 用CSS选择器定位Git stats模块下的strong标签 commit_strong = soup.select_one('h2.sr-only:contains("Git stats") + ul.list-style-none li strong') if commit_strong: commit_count = commit_strong.text.strip() print(f"提交次数:{commit_count}")
这个选择器明确指定了从Git stats标题开始,到后续列表里的strong元素,精准度很高。
总结
之前的代码不稳定就是因为选择器太宽泛,只要我们结合目标元素所在的专属模块上下文或者唯一属性标识,就能避免误匹配,稳定获取提交数啦。
内容的提问来源于stack exchange,提问作者user13295460
相关产品推荐
相关产品推荐

