Python使用Selenium接受Cookie后爬取网页仍无法获取正确HTML求助
问题解决方案
问题根因
- 未设置元素等待逻辑,Cookie同意按钮还未完全渲染到DOM中就执行点击操作,大概率导致点击未生效
- 点击Cookie按钮后,页面核心商品数据为异步渲染加载,直接同步获取
page_source只能拿到未完成内容渲染的初始页面代码 - 未做反爬规避配置,站点的反爬机制识别到Selenium自动化特征,返回了异常页面内容
修正后可运行代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # 替换为你本地的ChromeDriver实际路径 DRIVER_PATH = "你的ChromeDriver路径" url = 'https://www.wikiparfum.fr/explore/by-name?query=dior' # 配置Chrome参数,规避反爬检测 chrome_options = Options() # 移除Selenium自动化标识 chrome_options.add_argument('--disable-blink-features=AutomationControlled') chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(executable_path=DRIVER_PATH, options=chrome_options) # 修改navigator.webdriver属性,进一步规避反爬识别 driver.execute_cdp_cmd('Page.addScriptToEvaluateOnNewDocument', { 'source': 'Object.defineProperty(navigator, "webdriver", {get: () => undefined})' }) driver.get(url) # 设置最长15秒的显式等待 wait = WebDriverWait(driver, 15) # 等待Cookie按钮可点击后再执行点击操作 accept_btn = wait.until(EC.element_to_be_clickable((By.ID, 'onetrust-accept-btn-handler'))) accept_btn.click() # 等待页面核心香水卡片元素加载完成后,再获取页面源码 wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div[class*="perfume-card"]'))) html = driver.page_source soup = BeautifulSoup(html, 'lxml') print(soup) # 爬虫运行结束后关闭浏览器进程 driver.quit()
可选优化排查方向
- 如果访问站点网络延迟较高,可适当延长显式等待的超时时间
- 不需要可视化浏览器窗口时,可添加
--headless=new参数启用无头模式,同时也能降低被反爬识别的概率 - 若触发IP限流机制,可搭配代理IP服务使用
内容的提问来源于stack exchange,提问作者icarus-26
相关产品推荐
相关产品推荐

