使用BeautifulSoup4编写CNN新闻爬虫遇IndexError报错求助
问题分析与解决
报错原因
你遇到的IndexError: list index out of range是因为doc.find_all(text=f"{topic}")没有找到任何匹配内容,返回的prices是空列表,此时访问prices[0]自然触发索引越界错误。
导致查找失败的核心原因:
- 精确匹配限制:
find_all(text=...)默认做完全精确匹配,但CNN页面中的相关文本可能和你输入的"Coronavirus"不完全一致(比如大小写不同、带有空格/换行,或是页面内用小写"coronavirus")。 - 页面内容动态性:CNN首页内容实时更新,可能当前页面恰好没有包含该关键词的文本。
修复方案
方案1:正则表达式模糊匹配(忽略大小写)
用正则实现不区分大小写的模糊查找,扩大匹配范围:
from bs4 import BeautifulSoup import requests import re url = "https://www.cnn.com/" topic = input("What kind of news are you looking for? ") result = requests.get(url) doc = BeautifulSoup(result.text, "html.parser") # 匹配包含关键词的文本,忽略大小写 prices = doc.find_all(text=re.compile(topic, re.IGNORECASE)) if prices: parent = prices[0].parent print(parent) else: print(f"No news found related to {topic}")
方案2:先判断列表是否为空
在访问列表元素前先检查内容,避免索引错误:
from bs4 import BeautifulSoup import requests url = "https://www.cnn.com/" topic = input("What kind of news are you looking for? ") result = requests.get(url) doc = BeautifulSoup(result.text, "html.parser") prices = doc.find_all(text=f"{topic}") if len(prices) > 0: parent = prices[0].parent print(parent) else: print(f"No matching text found for {topic}")
方案3:定向查找含关键词的标签
如果新闻标题集中在特定标签(如<a>、<h3>),可以直接查找包含关键词的标签,更精准:
from bs4 import BeautifulSoup import requests url = "https://www.cnn.com/" topic = input("What kind of news are you looking for? ") result = requests.get(url) doc = BeautifulSoup(result.text, "html.parser") # 查找所有包含关键词的a标签(适配新闻标题常见载体) news_links = doc.find_all("a", string=lambda text: text and topic.lower() in text.lower()) if news_links: print(news_links[0]) else: print(f"No news links found related to {topic}")
内容的提问来源于stack exchange,提问作者yavda
相关产品推荐
相关产品推荐

