Python轮换User-Agent爬取Google Scholar仍无数据(Laravel调用)
Google Scholar爬取失败:AttributeError('NoneType'无find_all属性)
在Laravel项目中调用Python脚本爬取Google Scholar时,已配置随机轮换User-Agent,但始终无法获取目标数据,执行脚本时触发AttributeError,提示'NoneType' object has no attribute 'find_all',原因是通过soup.find('table',{'id':'gsc_a_t'})未找到目标表格元素,返回了None。
原Python脚本
import sys import pandas as pd from bs4 import BeautifulSoup import requests import random user_agent_list = [ 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_5) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/13.1.1 Safari/605.1.15', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:77.0) Gecko/20100101 Firefox/77.0', 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_5) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.97 Safari/537.36', 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:77.0) Gecko/20100101 Firefox/77.0', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/83.0.4103.97 Safari/537.36', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/74.0.3729.157 Safari/537.36', 'Mozilla/5.0 (Windows NT 5.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/46.0.2490.71 Safari/537.36', 'Opera/9.80 (Windows NT 6.1; WOW64) Presto/2.12.388 Version/12.18', 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.82 Safari/537.36 OPR/84.0.4316.14', 'Opera/9.80 (Linux armv7l) Presto/2.12.407 Version/12.51 , D50u-D1-UHD/V1.5.16-UHD (Vizio, D50u-D1, Wireless)', 'Opera/9.80 (Windows NT 6.0) Presto/2.12.388 Version/12.14', 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/55.0.2883.87 Safari/537.36 OPR/42.0.2393.94', 'Mozilla/5.0 (Linux; U; Android 8.1.0; zh-CN; EML-AL00 Build/HUAWEIEML-AL00) AppleWebKit/537.36 (KHTML, like Gecko) Version/4.0 Chrome/57.0.2987.108 baidu.sogo.uc.UCBrowser/11.9.4.974 UWS/2.13.1.48 Mobile Safari/537.36 AliApp(DingTalk/4.5.11) com.alibaba.android.rimet/10487439 Channel/227200 language/zh-CN', 'Mozilla/5.0 (X11; U; Linux i686; en-US) U2/1.0.0 UCBrowser/9.3.1.344', 'Mozilla/5.0 (Linux; U; Android 10; en-US; RMX1901 Build/QKQ1.190918.001) AppleWebKit/537.36 (KHTML, like Gecko) Version/4.0 Chrome/78.0.3904.108 UCBrowser/13.4.0.1306 Mobile Safari/537.36', 'UCWEB/2.0 (Java; U; MIDP-2.0; Nokia203/20.37) U2/1.0.0 UCBrowser/8.7.0.218 U2/1.0.0 Mobile' ] for _ in user_agent_list: #Pick a random user agent user_agent = random.choice(user_agent_list) #Set the headers headers = {'User-Agent': user_agent} url = 'https://scholar.google.com/citations?user=EnegzCwAAAAJ&hl=&view_op=list_works&cstart=0&pagesize=100' response=requests.get(url,headers=headers) soup=BeautifulSoup(response.text,'lxml') table = soup.find('table',{'id':'gsc_a_t'}) titles = [] for item in table.find_all(class_='gsc_a_at'): title = item.text titles.append(title) print(titles)
错误信息
Traceback (most recent call last): File "D:\Research SPIT\scraper\public\title.py", line 44, in
for item in table.find_all(class_='gsc_a_at'): AttributeError: 'NoneType' object has no attribute 'find_all'
Laravel控制器代码
$process = new Process(['python', public_path() . '/title.py'], null, ['SYSTEMROOT' => getenv('SYSTEMROOT'), 'PATH' => getenv("PATH")]); $process->run(); if (!$process->isSuccessful()) { dd($process->getErrorOutput()); } dd($process->getOutput());
问题原因
- User-Agent逻辑冗余:原代码循环遍历
user_agent_list反复随机选择User-Agent,最终仅使用最后一次循环的结果,逻辑无效且无意义。 - URL格式错误:URL中的
&是HTML转义字符,直接使用会导致请求地址不正确,无法访问目标页面。 - 反爬机制拦截:Google Scholar的反爬系统会检测请求的IP、会话、请求频率等特征,仅靠随机User-Agent无法绕过,会返回无目标内容的页面,导致找不到指定表格。
解决方案
1. 修正User-Agent逻辑
删除多余的循环,直接随机选择一次User-Agent:
# 替换原循环代码块 user_agent = random.choice(user_agent_list) headers = {'User-Agent': user_agent}
2. 修复URL格式
将URL中的&替换为&,还原正确的请求地址:
url = 'https://scholar.google.com/citations?user=EnegzCwAAAAJ&hl=&view_op=list_works&cstart=0&pagesize=100'
3. 增加反爬绕过措施
维持会话状态
使用requests.Session()保持会话,模拟正常用户的访问流程:
session = requests.Session() session.headers.update({'User-Agent': user_agent}) response = session.get(url)
添加请求延迟
随机延迟请求,避免因频率过高触发反爬:
import time time.sleep(random.uniform(1, 3)) # 请求前随机延迟1-3秒 response = requests.get(url, headers=headers)
使用代理IP
如果当前IP被封禁,可配置代理IP更换请求来源:
proxies = { 'http': 'http://your-proxy-ip:port', 'https': 'https://your-proxy-ip:port' } response = requests.get(url, headers=headers, proxies=proxies)
4. 添加异常处理
避免因找不到元素导致脚本崩溃,同时可排查是否被反爬拦截:
table = soup.find('table',{'id':'gsc_a_t'}) titles = [] if table: for item in table.find_all(class_='gsc_a_at'): title = item.text titles.append(title) else: print("未找到目标表格,可能被反爬拦截") # 打印返回页面内容,确认是否为拦截页面 # print(response.text)
5. 可选:使用无头浏览器渲染页面
若Google Scholar需要JS渲染才能加载内容,可使用selenium配合无头浏览器:
from selenium import webdriver from selenium.webdriver.chrome.options import Options chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument(f'user-agent={user_agent}') driver = webdriver.Chrome(options=chrome_options) driver.get(url) soup = BeautifulSoup(driver.page_source, 'lxml') table = soup.find('table',{'id':'gsc_a_t'}) # 后续处理逻辑同前 titles = [] if table: for item in table.find_all(class_='gsc_a_at'): titles.append(item.text) driver.quit() print(titles)
内容的提问来源于stack exchange,提问作者Ana Fernandes
相关产品推荐
相关产品推荐

