You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+BeautifulSoup爬取Google Scholar检索结果总条数

解决方案

方法1:BS4直接定位标签

Google Scholar页面的总检索结果数默认放在顶部class为gs_ab_md的<div>标签中,你使用的是意大利语界面,该标签的文本内容展示格式为Circa X risultati,其中X就是总结果数。
你可以用如下代码实现提取:

import re
from bs4 import BeautifulSoup
import requests

# 必须加User-Agent否则会被反爬拦截返回异常页面
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
url = "你的检索链接"
resp = requests.get(url, headers=headers)
soup = BeautifulSoup(resp.text, "lxml")

# 定位目标标签,取第一个符合条件的即可
result_stat = soup.find("div", class_="gs_ab_md")
if result_stat:
    stat_text = result_stat.get_text()
    # 意大利语千分位为点分隔,适配格式提取数字
    match_res = re.search(r"Circa ([\d\.]+) risultati", stat_text)
    if match_res:
        total_count = int(match_res.group(1).replace(".", ""))
        print(f"总结果数:{total_count}")

方法2:全页正则匹配(兼容性更强)

如果页面结构有变动导致标签定位失败,可以直接在请求返回的完整HTML文本里用正则匹配结果数,不需要依赖bs4的标签定位:

# 接上面的请求代码
match_res = re.search(r"Circa ([\d\.]+) risultati", resp.text)
if match_res:
    total_count = int(match_res.group(1).replace(".", ""))
    print(f"总结果数:{total_count}")

注意事项

  • Google Scholar有严格的反爬机制,请求频率不要太高,最好每次请求间隔2-3秒,避免IP被临时封禁
  • 必须携带合理的User-Agent请求头,否则返回的是反爬拦截页,找不到目标内容

内容的提问来源于stack exchange,提问作者Giorgia Mancini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 23:36:05