You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+Selenium谷歌购物爬虫本地正常,Ubuntu服务器运行异常

Google Shopping爬虫本地正常但Ubuntu服务器无数据的排查与修复

核心问题

本地运行正常的Google Shopping爬虫,部署到Ubuntu服务器后返回There is nothing to add,未抓取到任何数据,本质是服务器环境与本地环境的差异导致页面渲染、元素定位或驱动运行异常。

逐一排查与修复步骤

1. ChromeDriver路径与环境适配问题

你的代码中硬编码了Mac本地的ChromeDriver路径:

driver = webdriver.Chrome(options=options, executable_path="/Users/kevin/Documents/projects/deal_hunt/scraper_scripts/chromedriver")

Ubuntu服务器不存在该路径,且需要匹配服务器上Chrome浏览器的版本。

解决方法:

  • 先安装Ubuntu版Chrome浏览器:
    sudo apt update && sudo apt install chromium-browser
    
  • 安装对应版本的ChromeDriver,或用webdriver-manager自动管理(推荐):
    1. 安装依赖:pip install webdriver-manager
    2. 修改驱动初始化代码:
      from selenium.webdriver.chrome.service import Service
      from webdriver_manager.chrome import ChromeDriverManager
      
      options = Options()
      # 后续添加headless等参数
      driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
      

2. Headless模式与服务器无桌面环境适配

本地开启了可视化模式(options.headless = False),但Ubuntu服务器通常无桌面环境,直接运行会导致Chrome无法启动或渲染异常。

解决方法:
修改ChromeOptions,添加服务器运行必需的参数:

options = Options()
options.headless = True
options.add_argument('--no-sandbox')  # 绕过Ubuntu的安全沙箱限制
options.add_argument('--disable-dev-shm-usage')  # 解决服务器临时内存不足问题
options.add_argument('--window-size=1920,1080')  # 设置标准窗口大小,避免元素布局异常
options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36')  # 模拟正常浏览器UA,避免被Google拦截

3. 页面加载不完整导致元素定位失败

服务器网络延迟或Google反爬机制,可能导致driver.get(url)后页面未完全加载就提取源码,导致BeautifulSoup找不到目标元素(如i0X6df)。

解决方法:
替换固定等待为显式等待,确保页面核心元素加载完成后再提取源码:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def get_items(url, category):
    driver.get(url)
    # 等待搜索结果容器加载,超时10秒
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "i0X6df"))
    )
    content = driver.page_source
    soup = BeautifulSoup(content, features="lxml")
    # 后续代码...

4. 动态元素定位失效

Google Shopping的class名称(如i0X6df、_-oQ)可能是动态生成的,服务器环境下的页面结构与本地存在差异,导致定位失败。

解决方法:

  • 在服务器上保存页面源码,对比本地结构:
    在content = driver.page_source后添加:
    with open('server_page.html', 'w', encoding='utf-8') as f:
        f.write(content)
    
    下载server_page.html后,检查目标元素的实际class或XPath,替换为更稳定的定位方式(如XPath包含固定文本,或部分匹配class)。

5. 权限与依赖缺失

  • 确保ChromeDriver有可执行权限:chmod +x /path/to/chromedriver
  • 安装所有依赖库:pip install selenium beautifulsoup4 lxml

验证修改

修改完成后,在Ubuntu服务器上重新运行python3 main.py,若仍有问题,可逐步打印中间变量(如soup.find_all(attrs="i0X6df")的长度、items的长度),定位具体哪一步失败。

内容的提问来源于stack exchange,提问作者Kevin Ilondo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 02:27:46