使用scholarly调用Google Scholar报错:Cannot Fetch from Google Scholar
Google Scholar 数据获取报错排查(MaxTriesExceededException)
问题重现
使用scholarly 1.7.6版本,已开启VPN,运行以下代码:
from scholarly import scholarly search_query = scholarly.search_author('Albert Einstein') author = next(search_query).fill()
出现报错:
MaxTriesExceededException: Cannot Fetch from Google Scholar.
可能的原因及解决方向
- 反爬机制拦截:Google Scholar对非官方请求识别严格,旧版本scholarly的请求头模拟不够贴近真实浏览器,容易被判定为爬虫。可以尝试给scholarly配置自定义请求头,添加真实浏览器的User-Agent等信息。
- VPN节点IP被拉黑:当前所用VPN节点的IP可能因过往爬取行为被Google标记,换几个不同地区的VPN节点再测试。
- 版本过时:1.7.6版本较旧,Google Scholar页面结构可能已更新,导致旧版本的解析逻辑失效。执行
pip install --upgrade scholarly升级到最新版本重试。 - 请求频率超限:短时间内多次运行请求触发了Google的频率限制,暂停一段时间再尝试,或在代码中加入
time.sleep()设置请求间隔。 - 依赖库不兼容:检查
requests、beautifulsoup4等依赖库的版本是否与scholarly 1.7.6适配,必要时重新安装依赖:pip install --force-reinstall scholarly==1.7.6。
内容的提问来源于stack exchange,提问作者Mateo
相关产品推荐
相关产品推荐

