You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬Hacker News时遇IndexError:列表索引越界如何修复?

爬取Hacker News时触发IndexError的解决方法

尝试使用Python的requests和BeautifulSoup库爬取Hacker News的新闻标题,运行代码时出现索引越界错误。

原代码

import requests
from bs4 import BeautifulSoup

res = requests.get('https://news.ycombinator.com/news')
soup = BeautifulSoup(res.text, 'html.parser')
links = soup.select('.titleline a')
subtext = soup.select('.subtext')


def create_custom_hn(links, subtext):
    hn = []
    for index, item in enumerate(links):
        title = links[index].getText()
        href = links[index].get('href', None)
        votes = subtext[index].select('.score')
        if len(votes):
            points = int(votes[0].getText().replace(' points', ''))
            print(points)
            hn.append({'title': title, 'href': href})
    return hn


print(create_custom_hn(links, subtext))

错误信息

votes = subtext[index].select('.score')
            ~~~~~~~^^^^^^^
IndexError: list index out of range

问题根源

links和subtext两个列表长度不一致。Hacker News页面中,部分条目(比如置顶的推广内容、招聘广告)只有.titleline元素,没有对应的.subtext区域,导致subtext的长度比links短,循环到后续索引时,subtext[index]就会超出范围。

修复方案

不要通过索引强行关联两个独立列表,而是直接遍历页面中包含标题和子文本的完整新闻条目容器。Hacker News的每条新闻都包裹在tr.athing标签内,我们可以先抓取所有完整条目,再从每个条目里提取对应内容,确保一一对应。

修复后的代码:

import requests
from bs4 import BeautifulSoup

res = requests.get('https://news.ycombinator.com/news')
soup = BeautifulSoup(res.text, 'html.parser')
# 抓取所有完整的新闻条目容器
items = soup.select('.athing')

def create_custom_hn(items):
    hn = []
    for item in items:
        # 提取当前条目的标题和链接
        title_line = item.select_one('.titleline a')
        if not title_line:
            continue
        title = title_line.getText()
        href = title_line.get('href', None)
        
        # 找到当前条目对应的子文本行(下一个兄弟tr元素)
        subtext_row = item.find_next_sibling('tr')
        if not subtext_row:
            continue
        subtext = subtext_row.select_one('.subtext')
        if not subtext:
            continue
            
        votes = subtext.select('.score')
        if votes:
            points = int(votes[0].getText().replace(' points', ''))
            hn.append({'title': title, 'href': href, 'points': points})
    return hn

print(create_custom_hn(items))

关键改进点

  1. 以.athing为单位遍历,保证标题和子文本的对应关系
  2. 用find_next_sibling('tr')定位对应子文本行,避免索引不匹配
  3. 增加空值判断,防止异常条目导致代码崩溃

内容的提问来源于stack exchange,提问作者Hetarth7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 19:40:57