如何用Python+BeautifulSoup爬取Google Scholar检索结果总条数
解决方案
方法1:BS4直接定位标签
Google Scholar页面的总检索结果数默认放在顶部class为gs_ab_md的<div>标签中,你使用的是意大利语界面,该标签的文本内容展示格式为Circa X risultati,其中X就是总结果数。
你可以用如下代码实现提取:
import re from bs4 import BeautifulSoup import requests # 必须加User-Agent否则会被反爬拦截返回异常页面 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } url = "你的检索链接" resp = requests.get(url, headers=headers) soup = BeautifulSoup(resp.text, "lxml") # 定位目标标签,取第一个符合条件的即可 result_stat = soup.find("div", class_="gs_ab_md") if result_stat: stat_text = result_stat.get_text() # 意大利语千分位为点分隔,适配格式提取数字 match_res = re.search(r"Circa ([\d\.]+) risultati", stat_text) if match_res: total_count = int(match_res.group(1).replace(".", "")) print(f"总结果数:{total_count}")
方法2:全页正则匹配(兼容性更强)
如果页面结构有变动导致标签定位失败,可以直接在请求返回的完整HTML文本里用正则匹配结果数,不需要依赖bs4的标签定位:
# 接上面的请求代码 match_res = re.search(r"Circa ([\d\.]+) risultati", resp.text) if match_res: total_count = int(match_res.group(1).replace(".", "")) print(f"总结果数:{total_count}")
注意事项
- Google Scholar有严格的反爬机制,请求频率不要太高,最好每次请求间隔2-3秒,避免IP被临时封禁
- 必须携带合理的User-Agent请求头,否则返回的是反爬拦截页,找不到目标内容
内容的提问来源于stack exchange,提问作者Giorgia Mancini
相关产品推荐
相关产品推荐

