You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取Yahoo Finance新闻时遇IndexError:列表索引越界问题求助

解决Yahoo Finance新闻爬取的IndexError问题

问题根源

你的代码仅匹配了Ov(h) Pend(44px) Pstart(25px)这一种class的新闻容器,但页面实际存在另一种Ov(h) Pend(14%) Pend(44px)--sm1024的容器,导致爬取到的新闻数量不足20条,循环到第9次时触发索引越界错误。

解决方案

同时匹配两种class的容器,遍历所有找到的结果,避免固定次数循环的限制:

# 获取两种class的所有新闻容器
news_containers = soup3.find_all('div', class_=['Ov(h) Pend(44px) Pstart(25px)', 'Ov(h) Pend(14%) Pend(44px)--sm1024'])

# 遍历所有容器
for idx, container in enumerate(news_containers, start=1):
    # 获取标题(取最后一个a标签的文本)
    headline = container.find_all('a')[-1].text.strip()
    # 获取描述(取最后一个p标签的文本)
    description = container.find_all('p')[-1].text.strip()
    
    print(f"{idx}) {headline}")
    print(description)
    print()

关键改进点

  • 多class匹配:通过class_参数传入列表,一次性获取两种样式的新闻容器,不会遗漏新闻。
  • 动态遍历:用enumerate遍历所有找到的容器,不再固定循环20次,彻底避免索引越界问题。
  • 文本清理:添加strip()去除文本前后空白,让输出更整洁。

可选异常处理

如果担心部分容器缺少a/p标签导致报错,可增加异常捕获:

news_containers = soup3.find_all('div', class_=['Ov(h) Pend(44px) Pstart(25px)', 'Ov(h) Pend(14%) Pend(44px)--sm1024'])

for idx, container in enumerate(news_containers, start=1):
    try:
        headline = container.find_all('a')[-1].text.strip()
        description = container.find_all('p')[-1].text.strip()
        print(f"{idx}) {headline}")
        print(description)
    except IndexError:
        print(f"{idx}) 该新闻容器格式异常,跳过")
    print()

内容的提问来源于stack exchange,提问作者momo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 11:18:28