You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu 18.04无头模式下Selenium无法通过XPath定位元素问题

Ubuntu服务器Selenium Headless Chrome无法定位元素+白屏问题解决指南

问题背景

我在本地Windows机器上能正常以Headless模式运行Python Selenium爬虫脚本,截图、元素定位都没问题,但部署到Ubuntu服务器并通过cron调用时,却遇到了NoSuchElementException,具体报错如下:

File "DomainScraper.py", line 30, in
login = driver.find_element_by_xpath('//a[@href="'+login_url+'"]').click()
...
selenium.common.exceptions.NoSuchElementException: Message: no such element: Unable to locate element: {"method":"xpath","selector":"//a[@href="https://account.domaintools.com/log-in/?r=https%3A%2F%2Freversewhois.domaintools.com%2F%3Frefine"]"}
(Session info: headless chrome=81.0.4044.17)

我的核心代码片段如下:

#imports...
options = Options()
options.add_argument('--incognito')
options.add_argument('--headless')
options.add_argument('--no-sandbox')
options.add_argument('--disable-extensions')
options.add_argument('--disable-infobars')
options.add_argument('--allow-running-insecure-content')
driver = webdriver.Chrome(options=options)
driver.delete_all_cookies()
driver.implicitly_wait(3)
url = "https://reversewhois.domaintools.com/?refine#q=%5B%5B%5B%22whois%22%2C%222%22%2C%22VerifiedID%40SG-Mandatory%22%5D%5D%5D"
driver.get(url)
driver.save_screenshot("sample.png")
login_url = 'https://account.domaintools.com/log-in/?r=https%3A%2F%2Freversewhois.domaintools.com%2F%3Frefine'
login = driver.find_element_by_xpath('//a[@href="'+login_url+'"]').click()
# 后续登录及数据爬取代码

我尝试过安装xvfb和pyvirtualdisplay并添加显示配置,但问题依旧,而且发现不用xvfb时截图是白屏,怀疑Selenium根本没正常加载目标URL,该怎么解决?


分步解决方案

1. 先搞定Chrome与ChromeDriver的版本兼容性

你的Chrome版本是81.x,这个版本太老旧了,很可能和现代网站的JS逻辑不兼容,导致页面加载异常。建议直接升级到最新稳定版:

  • 卸载旧版本:
    sudo apt remove --purge google-chrome-stable
    
  • 安装最新Chrome:
    wget -q -O - https://dl-ssl.google.com/linux/linux_signing_key.pub | sudo apt-key add -
    echo "deb [arch=amd64] http://dl.google.com/linux/chrome/deb/ stable main" | sudo tee /etc/apt/sources.list.d/google-chrome.list
    sudo apt update && sudo apt install google-chrome-stable
    
  • 下载对应版本的ChromeDriver(必须和Chrome版本完全匹配),放到/usr/local/bin/并赋予权限:
    # 替换为对应版本的下载链接,比如Chrome 118.x对应的ChromeDriver
    wget https://chromedriver.storage.googleapis.com/118.0.5993.70/chromedriver_linux64.zip
    unzip chromedriver_linux64.zip
    sudo mv chromedriver /usr/local/bin/
    sudo chmod +x /usr/local/bin/chromedriver
    

2. 优化Headless Chrome的启动参数

服务器环境下,需要补充几个关键参数来模拟正常浏览器环境,避免被网站拦截或渲染异常:

options = Options()
options.add_argument('--incognito')
options.add_argument('--headless=new')  # 新版Headless模式,行为更接近正常Chrome
options.add_argument('--no-sandbox')    # 绕过Chrome的沙箱限制(服务器非桌面环境必需)
options.add_argument('--disable-extensions')
options.add_argument('--disable-infobars')
options.add_argument('--allow-running-insecure-content')
options.add_argument('--disable-gpu')   # 服务器无GPU,必须禁用
options.add_argument('--window-size=1920,1080')  # 设置足够大的窗口,避免元素因响应式布局隐藏
options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')  # 伪装成正常桌面浏览器UA
options.add_argument('--disable-dev-shm-usage')  # 解决服务器/dev/shm空间不足导致的崩溃问题

3. 替换隐式等待为显式等待,确保元素加载完成

implicitly_wait是全局等待,但可靠性不如显式等待,尤其是动态加载的页面。把查找元素的逻辑改成这样:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

driver.get(url)
# 最多等待10秒,直到登录按钮出现
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.XPATH, f'//a[@href="{login_url}"]'))
)
# 确认元素存在后再点击
login_btn = driver.find_element(By.XPATH, f'//a[@href="{login_url}"]')
login_btn.click()

尽量避免用time.sleep(),显式等待会根据元素加载状态自动调整等待时间,更稳定。

4. 排查cron的环境问题

cron的运行环境和你登录服务器后的环境不一样,经常会因为路径、环境变量导致问题:

  • 指定ChromeDriver绝对路径:在初始化driver时明确写ChromeDriver的位置,避免cron找不到:
    driver = webdriver.Chrome(executable_path='/usr/local/bin/chromedriver', options=options)
    
  • 设置脚本工作目录:cron默认的工作目录是用户根目录,所以截图、CSV文件可能会保存到你意想不到的地方,在脚本开头加上:
    import os
    # 替换成你的脚本所在的绝对路径
    os.chdir('/home/ubuntu/your-scraper-directory')
    
  • 输出日志排查问题:修改cron任务,把输出日志保存下来,方便调试:
    # 比如每分钟运行一次脚本,日志保存到scraper.log
    * * * * * python3 /home/ubuntu/your-scraper-directory/DomainScraper.py >> /home/ubuntu/your-scraper-directory/scraper.log 2>&1
    
    查看日志用cat /home/ubuntu/your-scraper-directory/scraper.log,能看到脚本运行时的所有报错信息。

5. 正确使用xvfb(如果仍需要)

如果还是依赖xvfb,确保显示配置正确,比如设置更大的分辨率:

from pyvirtualdisplay import Display
# 设置和Chrome窗口匹配的分辨率
display = Display(visible=0, size=(1920, 1080))
display.start()
# 然后再初始化driver

或者直接用xvfb-run命令运行脚本,这样不需要在代码里加pyvirtualdisplay的逻辑:

xvfb-run -a python3 DomainScraper.py

在cron里也可以这样用:

* * * * * xvfb-run -a python3 /home/ubuntu/your-scraper-directory/DomainScraper.py >> /home/ubuntu/your-scraper-directory/scraper.log 2>&1

内容的提问来源于stack exchange,提问作者user12110431

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 08:17:42