You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium点击按钮爬取文章:数据集为空求助

问题排查与修复方案

核心问题分析

  • 点击元素定位错误:你用/html/body/section/section/div[1]/div[1]/div/div/div[3]/div[4]/a/svg/use定位的是图标内部的<use>标签,这不是可点击的交互元素,外层的<a>标签才是真正的"Continua a leggere"按钮,导致elements列表为空,循环根本没执行。
  • 绝对XPATH脆弱:全路径XPATH极易因页面微小变化失效,应该用相对定位+文本/属性匹配。
  • 数据提取逻辑错误:获取标题和URL时用固定XPATH,每次都会取第一个文章的内容,不是当前操作的文章。
  • 缺少翻页逻辑:原代码完全没处理下一页的爬取逻辑。
  • 等待策略不合理:重复设置隐式等待,不如显式等待精准可靠。

修复后的完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
import pandas as pd

# 初始化浏览器驱动
driver = webdriver.Chrome()
wait = WebDriverWait(driver, 10)  # 显式等待对象,超时10秒
base_url = "https://www.beniculturali.it/comunicati/notizie-del-giorno"
driver.get(base_url)

# 处理Cookie弹窗
try:
    cookie_btn = wait.until(EC.element_to_be_clickable((By.CLASS_NAME, "cookiesjsr-btn")))
    cookie_btn.click()
except:
    pass

data = []

while True:
    # 等待当前页面的文章列表加载完成
    articles = wait.until(EC.presence_of_all_elements_located((By.XPATH, "//div[contains(@class, 'card')]")))
    
    for article in articles:
        try:
            # 获取文章标题和URL
            title_elem = article.find_element(By.TAG_NAME, "h3").find_element(By.TAG_NAME, "a")
            article_title = title_elem.text.strip()
            article_url = title_elem.get_attribute("href")
            
            # 点击"Continua a leggere"按钮
            read_more_btn = article.find_element(By.XPATH, ".//a[contains(text(), 'Continua a leggere')]")
            driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", read_more_btn)
            time.sleep(0.5)
            read_more_btn.click()
            
            # 等待文章详情加载,提取内容和日期
            wait.until(EC.presence_of_element_located((By.TAG_NAME, "article")))
            article_content = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='content-text']"))).text.strip()
            article_date = wait.until(EC.presence_of_element_located((By.TAG_NAME, "time"))).text.strip()
            
            # 存入数据列表
            data.append({
                'Titolo': article_title,
                'Data': article_date,
                'URL': article_url,
                'Contenuto': article_content
            })
            
            # 返回列表页
            driver.back()
            # 等待列表页重新加载完成
            wait.until(EC.presence_of_all_elements_located((By.XPATH, "//div[contains(@class, 'card')]")))
            
        except Exception as e:
            print(f"处理文章出错: {str(e)}")
            # 出错后返回列表页,避免后续操作异常
            driver.back()
            continue
    
    # 处理翻页:查找下一页按钮
    try:
        next_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//a[contains(@aria-label, 'Pagina successiva')]")))
        # 检查按钮是否禁用(如果是最后一页,按钮会有disabled类)
        if 'disabled' in next_btn.get_attribute("class"):
            print("已爬取到最后一页")
            break
        driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", next_btn)
        time.sleep(0.5)
        next_btn.click()
        # 等待下一页加载
        wait.until(EC.staleness_of(articles[0]))  # 等待原列表元素失效,说明页面已刷新
    except Exception as e:
        print(f"翻页出错或已无下一页: {str(e)}")
        break

# 关闭浏览器
driver.quit()

# 生成DataFrame并打印
df = pd.DataFrame(data)
print(df)
# 可选:保存到CSV
# df.to_csv('beni_culturali_notizie.csv', index=False, encoding='utf-8-sig')

关键修改说明

  1. 修正点击元素定位:用".//a[contains(text(), 'Continua a leggere')]"直接定位到可点击的<a>标签,确保能找到有效元素。
  2. 改用相对定位:所有元素定位都基于文章卡片的上下文(article.find_element),避免绝对路径的脆弱性。
  3. 添加翻页逻辑:通过查找"Pagina successiva"按钮,判断是否还有下一页,循环爬取所有页面。
  4. 优化等待策略:使用显式等待WebDriverWait确保元素加载完成,比time.sleep更可靠,同时减少不必要的等待时间。
  5. 修复数据提取:从当前文章卡片中获取标题和URL,保证数据对应正确。

内容的提问来源于stack exchange,提问作者Roberto Artiaco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 17:53:22