升级Selenium Python至v4.8.0后headless设置弃用警告及问题求助
Selenium 4.8.0 无头模式问题排查与解决
问题描述
将Selenium Python升级到v4.8.0版本后,使用options.headless = True出现DeprecationWarning。按照官方建议改用add_argument('--headless')或add_argument('--headless=new')后,代码仍然无法正常运行。
原代码
import time import pandas as pd from selenium import webdriver from selenium.webdriver import Chrome from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By options = webdriver.ChromeOptions() options.headless = True options.page_load_strategy = 'none' chrome_path = 'chromedriver.exe' chrome_service = Service(chrome_path) driver = Chrome(options=options, service=chrome_service) driver.implicitly_wait(5) url= "https://hk.centanet.com/findproperty/list/transaction/%E6%84%89%E6%99%AF%E6%96%B0%E5%9F%8E_3-DMHSZHHRHD?q=TiDxvVGMUUeutVzA0g1JlQ" driver.get(url) time.sleep(10) contents = driver.find_element(By.CSS_SELECTOR,"div[class*='bx--structured-list-tbody']") properties = contents.find_elements(By.CSS_SELECTOR,"div[class*='bx--structured-list-row']") def extract_data(element): columns = element.find_elements(By.CSS_SELECTOR,"div[class*='bx--structured-list-td']") Date = columns[0].text Dev = columns[1].text Price = columns[3].text RiseBox = columns[4].text Area = columns[5].text return{ "Date": Date, "Development": Dev, "Consideration": Price, "Change": RiseBox, "Area": Area } data = [] for property in properties: extracted_data = extract_data(property) data.append(extracted_data) df = pd.DataFrame(data) df.to_csv("result.csv", index=False)
警告信息
Selenium.py:10: DeprecationWarning: headless property is deprecated, instead use add_argument('--headless') or add_argument('--headless=new') options.headless = True
排查与解决方向
- 版本匹配检查:
--headless=new是Chrome 112及以上版本才支持的参数,如果你的Chrome版本低于112,使用该参数会导致启动失败。此时要么升级Chrome到对应版本,要么改用旧版无头参数--headless。同时要确保ChromeDriver版本与Chrome版本完全匹配,版本不兼容会直接导致驱动无法启动。 - 彻底替换无头配置:必须完全删除
options.headless = True这行代码,只保留options.add_argument('--headless')或options.add_argument('--headless=new'),两种配置共存可能引发冲突。 - 添加反检测参数:部分网站会识别无头浏览器并限制访问,补充以下参数模拟正常浏览器环境:
options.add_argument('--disable-blink-features=AutomationControlled') options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36') options.add_argument('--window-size=1920,1080') options.add_argument('--no-sandbox') - 优化页面等待逻辑:你使用了
page_load_strategy = 'none',这种模式下浏览器不会等待页面完全加载,time.sleep(10)的固定休眠不可靠,建议改用显式等待确保元素加载完成:from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 替换原有的time.sleep(10) WebDriverWait(driver, 20).until( EC.presence_of_element_located((By.CSS_SELECTOR, "div[class*='bx--structured-list-tbody']")) ) - 简化驱动管理:Selenium 4.6.0及以上版本自带驱动自动管理功能,无需手动指定
chromedriver.exe路径,直接初始化即可:# 移除chrome_path和chrome_service相关代码 driver = Chrome(options=options)
内容的提问来源于stack exchange,提问作者YPT
相关产品推荐
相关产品推荐

