Linux服务器部署Selenium爬虫遇点击及定位问题求助
问题诊断与修复方案
一、先检查Linux/Docker中Chrome的运行状态
1. 版本匹配检查
在Docker容器内执行以下命令,确认Chrome与Chromedriver版本完全一致:
google-chrome --version chromedriver --version
版本不匹配会直接导致Selenium无法正常驱动浏览器,是最常见的基础问题。
2. 基础可用性测试
在容器内运行Chrome测试页面加载,验证浏览器本身能否正常工作:
google-chrome --headless --no-sandbox --disable-dev-shm-usage https://www.ozon.ru --dump-dom > test.html
查看生成的test.html内容,如果页面被Cloudflare拦截(出现人机验证、跳转页面),则反爬是核心问题;如果内容为空或报错,说明Chrome环境配置有问题。
3. 页面截图排查
在代码中添加截图代码,直观查看服务器上浏览器渲染的页面状态:
driver.get(f"https://www.ozon.ru/product/{user_prompt}") time.sleep(10) driver.save_screenshot('ozon_page.png')
将截图从Docker容器复制到本地查看,如果显示Cloudflare验证页面,即可确认反爬拦截。
二、针对Cloudflare反爬的修复方案
1. 强化浏览器特征伪装
更新ChromeOptions,添加更多模拟真实用户的参数:
options = webdriver.ChromeOptions() # 基础参数保留 options.add_argument('--no-sandbox') options.add_argument('--headless=new') # 改用新版headless,更接近真实浏览器 options.add_argument('--disable-dev-shm-usage') options.add_argument("--disable-blink-features=AutomationControlled") # 新增参数 options.add_argument("--window-size=1920,1080") # 设置真实窗口大小 options.add_argument("--disable-gpu") options.add_argument("--disable-extensions") options.add_argument("--ignore-certificate-errors") options.add_argument("--disable-web-security") options.add_argument("--allow-running-insecure-content") # 更新为最新的真实User-Agent(从你的浏览器复制) options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36") options.add_argument("accept=text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8") options.add_argument("accept-language=en-US,en;q=0.5") options.add_argument("upgrade-insecure-requests=1") options.add_argument("sec-fetch-dest=document") options.add_argument("sec-fetch-mode=navigate") options.add_argument("sec-fetch-site=none") options.add_argument("sec-fetch-user=?1")
2. 深度隐藏自动化特征
扩展CDP命令,彻底屏蔽Selenium的识别标记:
driver.execute_cdp_cmd("Page.addScriptToEvaluateOnNewDocument", { 'source': ''' // 隐藏webdriver属性 Object.defineProperty(navigator, 'webdriver', { get: () => undefined }); // 删除Selenium特征变量 delete window.cdc_adoQpoasnfa76pfcZLmcfl_Array; delete window.cdc_adoQpoasnfa76pfcZLmcfl_Promise; delete window.cdc_adoQpoasnfa76pfcZLmcfl_Symbol; // 模拟真实浏览器的navigator属性 Object.defineProperty(navigator, 'plugins', { get: () => [1, 2, 3, 4, 5] }); Object.defineProperty(navigator, 'languages', { get: () => ['en-US', 'en'] }); ''' })
3. 替换为Undetected Chromedriver
改用undetected-chromedriver库,它专门针对反爬优化,自动隐藏所有Selenium特征:
安装命令:
pip install undetected-chromedriver
替换初始化代码:
import undetected_chromedriver as uc options = uc.ChromeOptions() options.add_argument('--no-sandbox') options.add_argument('--headless=new') options.add_argument('--disable-dev-shm-usage') # 其他参数同上 driver = uc.Chrome(options=options)
4. 优化等待策略
彻底替换time.sleep为显式等待,避免因页面加载延迟导致元素定位失败:
from selenium.webdriver.support import expected_conditions as EC wait = WebDriverWait(driver, 20) # 最长等待20秒 # 等待按钮可点击后再点击 btn = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, 'pl6'))) btn.click() # 等待SKU元素加载完成后再获取文本 sku = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="layoutPage"]/div[1]/div[3]/div[2]/div/div/div[2]/span'))).text
三、代码核心问题修复
1. 修复缩进与语法错误
driver.maximize_window()缩进错误,需要与前面的代码对齐if name == "main"改为if __name__ == "__main__"- 变量重复定义:
brand被多次赋值,容易引发错误
2. 避免重复定位元素
循环中重复定位pl6按钮,改用已获取的元素列表:
main_div = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'pl6'))) for btn in main_div: btn.click() # 等待页面更新 wait.until(EC.staleness_of(btn)) # 等待旧元素失效,确保页面已刷新 # 后续操作...
3. 处理元素定位异常
对每个元素定位添加异常捕获,避免单次失败导致整个请求崩溃:
try: sku = wait.until(EC.presence_of_element_located((By.XPATH, '//*[@id="layoutPage"]/div[1]/div[3]/div[2]/div/div/div[2]/span'))).text except Exception: sku = "N/A"
四、Docker环境优化
- 运行容器时增加共享内存:
docker run --shm-size=2g your-image-name
避免因默认共享内存不足导致Chrome崩溃。
- 使用官方Selenium镜像作为基础镜像,确保Chrome与Chromedriver版本匹配:
FROM selenium/standalone-chrome:latest # 安装Python、依赖等
内容的提问来源于stack exchange,提问作者whizzkid
相关产品推荐
相关产品推荐

