Python Selenium爬取网页时Jupyter单元加载未完成求助
问题分析与解决方案
你的代码单元加载不完大概率是因为Cookie弹窗未处理,这类网站通常会在用户未接受Cookie时阻塞页面后续渲染或脚本执行,导致Selenium无法正常定位元素或页面一直处于加载状态。以下是针对性的修正方案:
1. 核心优化点
- 用显式等待替代固定
sleep:避免因网络延迟或页面加载慢导致元素未出现的问题 - 优先处理Cookie弹窗:确保页面后续内容能正常渲染
- 改用更稳定的元素定位逻辑:降低页面结构变化带来的定位失败风险
2. 修正后的完整代码
import re from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 配置Chrome浏览器选项 options = Options() options.add_argument('window-size=1000,800') # 可选:规避部分网站的自动化检测 options.add_argument('--disable-blink-features=AutomationControlled') # 初始化浏览器并访问目标页面 navegador = webdriver.Chrome(options=options) navegador.get('https://www.politicos.org.br/Ranking') try: # 等待Cookie接受按钮加载完成并点击(适配页面的"Aceitar todos"按钮) cookie_accept_btn = WebDriverWait(navegador, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "Aceitar todos")]')) ) cookie_accept_btn.click() # 等待目标按钮加载完成并点击 target_btn = WebDriverWait(navegador, 10).until( EC.element_to_be_clickable((By.XPATH, '//*[@id="__next"]/div[2]/div[1]/div[4]/button')) ) target_btn.click() except Exception as e: print(f"操作出错: {str(e)}") finally: # 按需选择是否关闭浏览器 # navegador.quit() pass
额外排查要点
- 确认ChromeDriver版本与本地Chrome浏览器版本完全匹配,版本不兼容会导致页面加载异常
- 如果页面仍卡顿,可尝试添加
options.add_argument('--headless=new')启用无头模式测试
内容的提问来源于stack exchange,提问作者Andre Luiz Moura
相关产品推荐
相关产品推荐

