爬取Yahoo Finance新闻时遇IndexError:列表索引越界问题求助
解决Yahoo Finance新闻爬取的IndexError问题
问题根源
你的代码仅匹配了Ov(h) Pend(44px) Pstart(25px)这一种class的新闻容器,但页面实际存在另一种Ov(h) Pend(14%) Pend(44px)--sm1024的容器,导致爬取到的新闻数量不足20条,循环到第9次时触发索引越界错误。
解决方案
同时匹配两种class的容器,遍历所有找到的结果,避免固定次数循环的限制:
# 获取两种class的所有新闻容器 news_containers = soup3.find_all('div', class_=['Ov(h) Pend(44px) Pstart(25px)', 'Ov(h) Pend(14%) Pend(44px)--sm1024']) # 遍历所有容器 for idx, container in enumerate(news_containers, start=1): # 获取标题(取最后一个a标签的文本) headline = container.find_all('a')[-1].text.strip() # 获取描述(取最后一个p标签的文本) description = container.find_all('p')[-1].text.strip() print(f"{idx}) {headline}") print(description) print()
关键改进点
- 多class匹配:通过
class_参数传入列表,一次性获取两种样式的新闻容器,不会遗漏新闻。 - 动态遍历:用
enumerate遍历所有找到的容器,不再固定循环20次,彻底避免索引越界问题。 - 文本清理:添加
strip()去除文本前后空白,让输出更整洁。
可选异常处理
如果担心部分容器缺少a/p标签导致报错,可增加异常捕获:
news_containers = soup3.find_all('div', class_=['Ov(h) Pend(44px) Pstart(25px)', 'Ov(h) Pend(14%) Pend(44px)--sm1024']) for idx, container in enumerate(news_containers, start=1): try: headline = container.find_all('a')[-1].text.strip() description = container.find_all('p')[-1].text.strip() print(f"{idx}) {headline}") print(description) except IndexError: print(f"{idx}) 该新闻容器格式异常,跳过") print()
内容的提问来源于stack exchange,提问作者momo
相关产品推荐
相关产品推荐

