使用Selenium爬取麦肯锡网页时出现'NoneType'无'suppress'属性错误求助
解决Selenium爬取麦肯锡官网时的Service销毁错误
这个错误是Selenium驱动服务对象销毁时出现的资源泄漏问题,通常和驱动初始化方式、版本不匹配或未正确关闭浏览器有关,以下是具体解决步骤:
1. 匹配Selenium与浏览器驱动版本
无论使用Chrome还是Firefox,必须保证webdriver(ChromeDriver/GeckoDriver)版本和浏览器版本完全对应。嫌手动下载麻烦的话,可以用webdriver-manager库自动管理驱动,安装命令:
pip install webdriver-manager
2. 修正驱动初始化代码
新版本Selenium推荐用Service类管理驱动,替代旧的直接传路径写法,避免服务初始化异常:
- Chrome示例:
from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager # 自动下载并匹配对应版本的ChromeDriver browser = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=chrome_options)
- Firefox示例:
from selenium.webdriver.firefox.service import Service from webdriver_manager.firefox import GeckoDriverManager browser = webdriver.Firefox(service=Service(GeckoDriverManager().install()))
3. 代码末尾正确关闭浏览器
爬取完成后必须调用browser.quit()彻底终止驱动服务,避免资源泄漏导致销毁时的错误:
# 在soup初始化后添加 browser.quit()
4. 优化元素定位方式
你原代码中加载更多按钮的XPATH依赖固定section索引(section[11]),页面结构更新就会定位失败,换成更稳定的定位方式,比如通过按钮class或文本:
# 替换原来的按钮定位代码 button = browser.find_element(By.XPATH, '//a[contains(@class, "load-more-button") or contains(text(), "Load more")]')
修改后的完整代码示例
from selenium.webdriver.common.by import By from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from bs4 import BeautifulSoup import time chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--disable-notifications") # 自动管理ChromeDriver browser = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=chrome_options) url = "https://www.mckinsey.com/capabilities/operations/our-insights" browser.get(url) time.sleep(5) try: accept = browser.find_element(By.XPATH, '//*[@id="onetrust-accept-btn-handler"]') accept.click() time.sleep(2) browser.execute_script("window.scrollTo(0, document.body.scrollHeight);") except Exception as e: print(f"处理cookie弹窗失败: {e}") n = 1 while n < 3: try: browser.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 更稳定的按钮定位 button = browser.find_element(By.XPATH, '//a[contains(@class, "load-more-button")]') button.click() time.sleep(2) browser.execute_script("window.scrollTo(0, document.body.scrollHeight);") print('page', n) n += 1 except Exception as e: print(f'page ended at {n}, 错误: {e}') break source = browser.execute_script("return document.body.innerHTML") time.sleep(5) soup = BeautifulSoup(source, 'lxml') # 必须关闭浏览器 browser.quit()
内容的提问来源于stack exchange,提问作者Vinay
相关产品推荐
相关产品推荐

