You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取infomoney.com.br遇ElementClickInterceptedException问题

解决Infomoney新闻爬取中“Carregar Mais”按钮点击拦截问题

问题说明

尝试爬取https://www.infomoney.com.br/ultimas-noticias/的新闻标题与链接时,点击底部蓝色“Carregar Mais”按钮加载更多内容,始终触发ElementClickInterceptedException错误,已尝试多种方案仍无效,原代码如下:

options = webdriver.ChromeOptions()
# options.add_argument("headless")  # 无头模式(已注释)
driver = webdriver.Chrome(options=options)

driver.get(jornal_infomoney)
time.sleep(40)  # 强制等待页面加载

for _ in range(5):
    driver.maximize_window()
    driver.implicitly_wait(10)
    button = driver.find_element(By.XPATH, "//div[@id='infinite-handle']//button")
    button.click()
    # button = driver.find_element(By.XPATH, "//button[text()='Carregar mais']")  # 通过按钮文本定位
    # driver.execute_script("arguments[0].scrollIntoView();", button)  # 滚动到按钮可见
    # ActionChains(driver).move_to_element(button).perform()  # 鼠标移动到按钮上
    # button = WebDriverWait(driver, 30).until(EC.element_to_be_clickable((By.XPATH, "//button[text()='Carregar mais']")))  # 等待按钮可点击
    # driver.execute_script("arguments[0].click();", button)  # 用JS执行点击
    '''
    WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, "//button[text()='Carregar mais']")))
    button.click()
    wait = WebDriverWait(driver, 10)
    load_more_button = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[text()='Carregar mais']")))
    load_more_button.click().perform()
    '''
    time.sleep(2)  # 点击后等待加载

results_infomoney = pd.DataFrame(columns=['link', 'titulo', 'data', 'conteudo'])
article_id = 0

soup = BeautifulSoup(driver.page_source, "html.parser")
articles = soup.find_all("div", class_="row py-3 item")

for article in articles:
    article_title = article.find("span", class_="hl-title hl-title-2").find("a")
    article_link = article_title["href"]
    new_article = pd.DataFrame([{'link':article_link, 'titulo':article_title}])
    results_infomoney = pd.concat([results_infomoney, new_article], ignore_index = True)
    article_id+=1

driver.quit()

results_infomoney.to_excel((output_folder + '/results_infomoney_v1.xlsx'))

解决方案

错误核心是按钮被页面元素(如Cookie弹窗、固定导航栏)遮挡,或未完全加载完成。以下是针对性优化方案:

  1. 先处理页面弹窗:网站加载时大概率会弹出Cookie授权窗口,需优先关闭避免遮挡按钮
  2. 精准等待元素状态:用WebDriverWait替代固定时长的time.sleep,确保按钮完全可交互
  3. 优化滚动定位:将按钮滚动到视窗中间,避免被顶部/底部固定栏遮挡
  4. JS强制点击:直接通过JavaScript触发按钮点击,绕过Selenium的点击拦截检测

修改后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import pandas as pd
import time

jornal_infomoney = "https://www.infomoney.com.br/ultimas-noticias/"
output_folder = "./"  # 替换为你的实际输出路径

options = webdriver.ChromeOptions()
# options.add_argument("headless")  # 如需无头模式,取消注释并添加--disable-gpu参数
driver = webdriver.Chrome(options=options)
driver.maximize_window()

driver.get(jornal_infomoney)

# 处理Cookie授权弹窗(如果存在)
try:
    cookie_btn = WebDriverWait(driver, 15).until(
        EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Aceitar todos')]"))
    )
    cookie_btn.click()
except:
    pass  # 无弹窗则跳过

# 循环加载更多内容
for _ in range(5):
    try:
        # 等待按钮加载完成
        load_btn = WebDriverWait(driver, 20).until(
            EC.presence_of_element_located((By.XPATH, "//div[@id='infinite-handle']//button[text()='Carregar mais']"))
        )
        # 滚动到按钮中间位置,避免被固定栏遮挡
        driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", load_btn)
        time.sleep(1)
        
        # 用JS执行点击,绕过遮挡问题
        driver.execute_script("arguments[0].click();", load_btn)
        
        # 等待新内容加载完成(检测旧元素失效)
        WebDriverWait(driver, 10).until(
            EC.staleness_of(driver.find_element(By.CLASS_NAME, "row.py-3.item"))
        )
    except Exception as e:
        print(f"加载更多失败: {str(e)}")
        break

# 解析并保存新闻数据
results_infomoney = pd.DataFrame(columns=['link', 'titulo', 'data', 'conteudo'])
soup = BeautifulSoup(driver.page_source, "html.parser")
articles = soup.find_all("div", class_="row py-3 item")

for article in articles:
    article_a = article.find("span", class_="hl-title hl-title-2").find("a")
    if article_a:
        article_link = article_a["href"]
        article_title = article_a.get_text(strip=True)
        new_row = pd.DataFrame([{'link': article_link, 'titulo': article_title}])
        results_infomoney = pd.concat([results_infomoney, new_row], ignore_index=True)

results_infomoney.to_excel(f"{output_folder}/results_infomoney_v1.xlsx", index=False)
driver.quit()

关键优化点

  • Cookie弹窗前置处理:消除最常见的遮挡源
  • 滚动定位优化:确保按钮处于视窗可交互区域
  • JS点击替代:绕过Selenium的点击拦截限制
  • 动态加载等待:用元素失效检测替代固定等待,更可靠

内容的提问来源于stack exchange,提问作者Carlos Eduardo Barros

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 05:23:13