You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium Python爬取金融稳定违约事件时间线页面报错求助:'WebDriver'对象无'find_elements_by_xpath'属性

Selenium Python爬取金融稳定违约事件时间线页面报错求助:'WebDriver'对象无'find_elements_by_xpath'属性

嗨,我来帮你搞定这个问题!你遇到的这个AttributeError是因为你用的是Selenium 4.x版本——这个版本已经彻底移除了find_elements_by_xpath、find_element_by_tag_name这类旧的元素定位方法,必须改用基于By类的新语法啦。

另外你提到页面可能是JavaScript heavy,意思是页面里的很多内容(比如时间线事件)是通过JavaScript动态渲染出来的,不是页面加载完就立刻显示的,只设置页面加载超时可能不够,得用显式等待来确保元素加载完成后再去定位,不然很容易找不到元素。

我帮你把代码修改好了,不仅解决了报错问题,还加上了日期提取(毕竟你需要事件和日期),同时处理了动态加载的情况:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 设置chromedriver路径(改成你本地的实际路径)
webdriver_service = Service('/Users/hazael/WEBSCRAPE/chromedriver')
driver = webdriver.Chrome(service=webdriver_service)

# 打开目标页面
url = 'https://carnegieendowment.org/specialprojects/protectingfinancialstability/timeline'
driver.get(url)

try:
    # 显式等待事件容器加载完成(最多等15秒),解决JS动态加载问题
    event_containers = WebDriverWait(driver, 15).until(
        EC.presence_of_all_elements_located((By.XPATH, "//div[@class='frst-timeline-content-inner']"))
    )

    # 提取事件名和日期
    for container in event_containers:
        # 用新语法获取事件名
        event_name = container.find_element(By.TAG_NAME, 'h2').text
        # 从时间线的时间标签提取日期
        event_date = container.find_element(By.XPATH, "./preceding-sibling::div[@class='frst-timeline-time']").text
        print(f"事件名称: {event_name}")
        print(f"事件日期: {event_date}")
        print("---")

    # 处理"Learn More"按钮(用显式等待确保按钮可点击)
    try:
        learn_more_button = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, "//div[@data-more-label='Learn More']"))
        )
        learn_more_button.click()
        print("成功点击Learn More按钮")
    except:
        print("Learn More按钮不可见或无法点击")

finally:
    # 操作完成后关闭浏览器,避免资源占用
    driver.quit()

关键修改点说明:

  • 把所有旧的find_elements_by_xpath、find_element_by_tag_name替换成find_elements(By.XPATH, ...)、find_element(By.TAG_NAME, ...)的新语法,适配Selenium 4.x版本
  • 引入WebDriverWait和expected_conditions做显式等待,确保动态加载的元素完全渲染后再操作,解决JS heavy页面的加载问题
  • 新增了日期提取逻辑,满足你需要获取事件和日期的核心需求
  • 用try...finally结构确保浏览器能正常关闭,避免残留进程

如果运行后还是有部分事件没加载出来,可能是页面需要滚动才能加载更多内容,到时可以再加一段滚动页面的逻辑哦~

备注:内容来源于stack exchange,提问作者Hazael Otuya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.22 10:58:14