You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将特定XPath转为通用路径批量提取雅虎财经页面所有新闻链接

问题原因

你当前使用的XPath中li[1]的下标是固定选取列表的第一个元素,所以只能获取到单篇新闻。同时多层级嵌套的路径写法容错性极低,页面结构稍有调整就会匹配失效。

修正后代码

import time
from selenium import webdriver
# 若使用Selenium 4.0+版本,取消注释下一行导入
# from selenium.webdriver.common.by import By

# 此处补充你的webdriver初始化逻辑,比如 driver = webdriver.Chrome()
driver.maximize_window()
driver.implicitly_wait(10)
driver.get("https://finance.yahoo.com/quote/AAPL/news?p=AAPL")
links = []
dates = []

# 滚动加载更多新闻
for i in range(20):
    driver.execute_script("window.scrollBy(0, 250)")
    time.sleep(1)

# 匹配所有新闻列表项,移除固定下标[1]
# Selenium 4.0+写法:all_items = driver.find_elements(By.XPATH, '//*[@id="latestQuoteNewsStream-0-Stream"]/ul/li')
all_items = driver.find_elements_by_xpath('//*[@id="latestQuoteNewsStream-0-Stream"]/ul/li')

for item in all_items:
    try:
        # 改用特征匹配,不写死层级,容错性更高
        news_a = item.find_element_by_xpath('.//h3/a[@href]')
        news_date = item.find_element_by_xpath('.//div[contains(@class,"byline")]//span[contains(@class,"date")]')
        # 要新闻标题就取news_a.text,要链接就取news_a.get_attribute('href')
        links.append(news_a.get_attribute('href'))
        dates.append(news_date.text)
    except Exception:
        # 跳过广告位、空占位项等非新闻节点
        continue

driver.quit()

核心修改说明

  • 移除XPath中li后的[1]固定下标,匹配列表内所有新闻对应的li节点
  • 放弃多层级嵌套的路径写法,改用特征标签模糊匹配,页面内部嵌套结构微调也不会失效
  • 增加异常捕获逻辑,自动跳过列表中的广告、无效占位项
  • 修正原代码中未导入time库、导入未使用的依赖的问题
  • 兼容Selenium新旧版本写法,可根据自己的环境选择对应调用方式

内容的提问来源于stack exchange,提问作者huy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 09:24:00