使用BeautifulSoup爬Hacker News时遇IndexError:列表索引越界如何修复?
爬取Hacker News时触发IndexError的解决方法
尝试使用Python的requests和BeautifulSoup库爬取Hacker News的新闻标题,运行代码时出现索引越界错误。
原代码
import requests from bs4 import BeautifulSoup res = requests.get('https://news.ycombinator.com/news') soup = BeautifulSoup(res.text, 'html.parser') links = soup.select('.titleline a') subtext = soup.select('.subtext') def create_custom_hn(links, subtext): hn = [] for index, item in enumerate(links): title = links[index].getText() href = links[index].get('href', None) votes = subtext[index].select('.score') if len(votes): points = int(votes[0].getText().replace(' points', '')) print(points) hn.append({'title': title, 'href': href}) return hn print(create_custom_hn(links, subtext))
错误信息
votes = subtext[index].select('.score') ~~~~~~~^^^^^^^ IndexError: list index out of range
问题根源
links和subtext两个列表长度不一致。Hacker News页面中,部分条目(比如置顶的推广内容、招聘广告)只有.titleline元素,没有对应的.subtext区域,导致subtext的长度比links短,循环到后续索引时,subtext[index]就会超出范围。
修复方案
不要通过索引强行关联两个独立列表,而是直接遍历页面中包含标题和子文本的完整新闻条目容器。Hacker News的每条新闻都包裹在tr.athing标签内,我们可以先抓取所有完整条目,再从每个条目里提取对应内容,确保一一对应。
修复后的代码:
import requests from bs4 import BeautifulSoup res = requests.get('https://news.ycombinator.com/news') soup = BeautifulSoup(res.text, 'html.parser') # 抓取所有完整的新闻条目容器 items = soup.select('.athing') def create_custom_hn(items): hn = [] for item in items: # 提取当前条目的标题和链接 title_line = item.select_one('.titleline a') if not title_line: continue title = title_line.getText() href = title_line.get('href', None) # 找到当前条目对应的子文本行(下一个兄弟tr元素) subtext_row = item.find_next_sibling('tr') if not subtext_row: continue subtext = subtext_row.select_one('.subtext') if not subtext: continue votes = subtext.select('.score') if votes: points = int(votes[0].getText().replace(' points', '')) hn.append({'title': title, 'href': href, 'points': points}) return hn print(create_custom_hn(items))
关键改进点
- 以
.athing为单位遍历,保证标题和子文本的对应关系 - 用
find_next_sibling('tr')定位对应子文本行,避免索引不匹配 - 增加空值判断,防止异常条目导致代码崩溃
内容的提问来源于stack exchange,提问作者Hetarth7
相关产品推荐
相关产品推荐

