You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium爬取网页时Jupyter单元加载未完成求助

问题分析与解决方案

你的代码单元加载不完大概率是因为Cookie弹窗未处理,这类网站通常会在用户未接受Cookie时阻塞页面后续渲染或脚本执行,导致Selenium无法正常定位元素或页面一直处于加载状态。以下是针对性的修正方案:

1. 核心优化点

  • 用显式等待替代固定sleep:避免因网络延迟或页面加载慢导致元素未出现的问题
  • 优先处理Cookie弹窗:确保页面后续内容能正常渲染
  • 改用更稳定的元素定位逻辑:降低页面结构变化带来的定位失败风险

2. 修正后的完整代码

import re
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 配置Chrome浏览器选项
options = Options()
options.add_argument('window-size=1000,800')
# 可选:规避部分网站的自动化检测
options.add_argument('--disable-blink-features=AutomationControlled')

# 初始化浏览器并访问目标页面
navegador = webdriver.Chrome(options=options)
navegador.get('https://www.politicos.org.br/Ranking')

try:
    # 等待Cookie接受按钮加载完成并点击(适配页面的"Aceitar todos"按钮)
    cookie_accept_btn = WebDriverWait(navegador, 10).until(
        EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "Aceitar todos")]'))
    )
    cookie_accept_btn.click()
    
    # 等待目标按钮加载完成并点击
    target_btn = WebDriverWait(navegador, 10).until(
        EC.element_to_be_clickable((By.XPATH, '//*[@id="__next"]/div[2]/div[1]/div[4]/button'))
    )
    target_btn.click()
except Exception as e:
    print(f"操作出错: {str(e)}")
finally:
    # 按需选择是否关闭浏览器
    # navegador.quit()
    pass

额外排查要点

  • 确认ChromeDriver版本与本地Chrome浏览器版本完全匹配,版本不兼容会导致页面加载异常
  • 如果页面仍卡顿,可尝试添加options.add_argument('--headless=new')启用无头模式测试

内容的提问来源于stack exchange,提问作者Andre Luiz Moura

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 13:25:16