You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用Selenium接受Cookie后爬取网页仍无法获取正确HTML求助

问题解决方案

问题根因

  • 未设置元素等待逻辑,Cookie同意按钮还未完全渲染到DOM中就执行点击操作,大概率导致点击未生效
  • 点击Cookie按钮后,页面核心商品数据为异步渲染加载,直接同步获取page_source只能拿到未完成内容渲染的初始页面代码
  • 未做反爬规避配置,站点的反爬机制识别到Selenium自动化特征,返回了异常页面内容

修正后可运行代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# 替换为你本地的ChromeDriver实际路径
DRIVER_PATH = "你的ChromeDriver路径"
url = 'https://www.wikiparfum.fr/explore/by-name?query=dior'

# 配置Chrome参数,规避反爬检测
chrome_options = Options()
# 移除Selenium自动化标识
chrome_options.add_argument('--disable-blink-features=AutomationControlled')
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)

driver = webdriver.Chrome(executable_path=DRIVER_PATH, options=chrome_options)
# 修改navigator.webdriver属性,进一步规避反爬识别
driver.execute_cdp_cmd('Page.addScriptToEvaluateOnNewDocument', {
    'source': 'Object.defineProperty(navigator, "webdriver", {get: () => undefined})'
})

driver.get(url)
# 设置最长15秒的显式等待
wait = WebDriverWait(driver, 15)

# 等待Cookie按钮可点击后再执行点击操作
accept_btn = wait.until(EC.element_to_be_clickable((By.ID, 'onetrust-accept-btn-handler')))
accept_btn.click()

# 等待页面核心香水卡片元素加载完成后,再获取页面源码
wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div[class*="perfume-card"]')))

html = driver.page_source
soup = BeautifulSoup(html, 'lxml')

print(soup)
# 爬虫运行结束后关闭浏览器进程
driver.quit()

可选优化排查方向

  • 如果访问站点网络延迟较高,可适当延长显式等待的超时时间
  • 不需要可视化浏览器窗口时,可添加--headless=new参数启用无头模式,同时也能降低被反爬识别的概率
  • 若触发IP限流机制,可搭配代理IP服务使用

内容的提问来源于stack exchange,提问作者icarus-26

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 06:36:03