You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Centos EC2实例中无头Chrome+Selenium爬取Everlane网站遭拦截引发NoSuchElementException问题求助

解决Selenium在EC2 CentOS上被Everlane反爬拦截的问题

我之前也遇到过类似的云服务器上Selenium被反爬拦截的情况,从你给出的页面返回内容(带Ray ID和浏览器检查提示)来看,Everlane用了Cloudflare的防护机制,加上你的自动化请求特征太明显,再配合EC2的 data center IP属性,很容易被标记为可疑流量。下面是几个实用的解决思路:

1. 强化ChromeOptions,模拟真实浏览器环境

原生headless Chrome的特征太容易被检测,你需要添加更多参数掩盖自动化痕迹,同时适配CentOS的运行环境:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

ser = Service(ChromeDriverManager().install())
chrome_options = webdriver.ChromeOptions()

# 核心配置:让请求更像真实用户
chrome_options.add_argument("--headless=new")  # 新版headless模式,行为更接近正常Chrome
chrome_options.add_argument("--disable-blink-features=AutomationControlled")  # 禁用自动化检测标记
chrome_options.add_argument("--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36")  # 模拟桌面浏览器UA
chrome_options.add_argument("--window-size=1920,1080")  # 设置正常桌面窗口尺寸
chrome_options.add_argument("--no-sandbox")  # CentOS环境必需,避免权限报错
chrome_options.add_argument("--disable-dev-shm-usage")  # 解决内存不足导致的进程崩溃
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])  # 隐藏"Chrome正在被自动化控制"提示
chrome_options.add_experimental_option('useAutomationExtension', False)  # 禁用默认的自动化扩展

selenium_driver = webdriver.Chrome(service=ser, options=chrome_options)

这些配置能大幅降低你的请求被识别为自动化工具的概率。

2. 用显式等待替代硬等待,适配页面加载流程

简单的sleep可能赶不上Cloudflare的跳转验证节奏,改用WebDriverWait等待目标元素出现,同时捕获超时状态方便排查:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException

url = 'https://www.everlane.com/products/womens-cloud-cable-knit-vest-oatmeal?collection=womens-newest-arrivals'
selenium_driver.get(url)

try:
    # 最长等待15秒,直到标题元素加载完成
    title = WebDriverWait(selenium_driver, 15).until(
        EC.presence_of_element_located((By.XPATH, '//*[@id="content"]/div/div[3]/div[2]/div/div/div/div[2]/div/div[1]/hgroup/h1/span'))
    )
    print(title.text)
except TimeoutException:
    # 超时后打印页面源码,确认是否仍卡在验证环节
    print("超时未找到目标元素,当前页面源码:")
    print(selenium_driver.page_source)

3. 改用undetected-chromedriver绕过检测

如果上面的配置还是无效,试试专门针对反爬优化的undetected-chromedriver库,它会自动修改ChromeDriver的底层特征,避开Cloudflare这类防护机制的检测:
先安装库:

pip install undetected-chromedriver

然后替换原有的webdriver初始化代码:

import undetected_chromedriver as uc

# 复用之前配置好的chrome_options
selenium_driver = uc.Chrome(options=chrome_options)

这个库处理了很多原生Selenium暴露的自动化特征,对Cloudflare的基础防护效果很显著。

4. 解决IP层面的拦截

EC2的IP属于数据中心IP段,很多网站会直接限制这类IP的访问。你可以尝试:

  • 更换EC2公网IP:停止EC2实例后重新启动,AWS会分配新的公网IP,有可能绕过IP黑名单。
  • 使用住宅代理:如果换IP无效,考虑使用住宅代理IP(而非数据中心代理),这类IP更接近真实用户的访问来源,不容易被拦截。

5. 检查服务器环境细节

  • 确保EC2实例的时间和时区与本地一致,很多网站会通过时间差判断请求是否异常。
  • 如果有条件,在EC2上安装桌面环境并手动访问目标网站,确认是否存在地区限制或需要手动完成的验证步骤。

建议先从修改ChromeOptions和试用undetected-chromedriver开始,这两个方案成本最低且见效快,如果还是不行再考虑更换IP或使用代理。

内容的提问来源于stack exchange,提问作者sm1994

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:50:10