You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python RSS抓取脚本长时间运行后报IndexError问题排查

异常触发原因

这个IndexError: list index out of range的直接触发原因是:程序执行到第16行访问f.entries[0]时,f.entries是空列表,索引0不存在,触发越界错误。

程序连续运行11小时才报错,和初始逻辑无关,是单次拉取RSS源时出现了临时异常,导致feedparser没有解析出任何文章条目,常见触发场景包括:

  • 树莓派当时出现临时网络波动、断网、连接RSS服务器超时,或被对方服务临时限流,没有拿到完整的RSS响应内容
  • Tagesschau网站当时临时维护、RSS接口短暂返回5xx/4xx错误页、空响应,返回内容不是标准RSS结构,无法解析出文章列表
  • 本地DNS解析临时故障,请求被导向错误地址,拿到的响应内容不包含文章条目
原代码的缺陷
  • 没有做边界校验:默认每次请求都能拿到合法的、包含至少1条文章的RSS内容,完全没处理空响应、解析失败的场景
  • 没有异常捕获机制:单次请求、解析失败会直接终止整个无限循环,程序直接退出,无法满足常驻运行的需求
  • 存在冗余代码:使用with open()上下文管理器操作文件时,会自动完成文件关闭,不需要额外手动调用myfile.close()
  • 打开文件未指定编码:写入德语特殊字符(ä/ö/ü/ß等)时,可能在部分系统环境下触发编码错误
修复后可稳定运行的代码
import feedparser
import datetime
import time

last_item = ""
x = 1
news_nr = 1
articles_found = 0
RSS_ADDRESS = "https://www.tagesschau.de/xml/rss2/"

while True:
    todays_date = datetime.date.today()
    today_formatted = todays_date.strftime("%d.%m.%Y")
    time_now = datetime.datetime.now()
    time_formatted = time_now.strftime("%H:%M:%S")
    
    try:
        f = feedparser.parse(RSS_ADDRESS)
        # 先校验条目列表非空,再访问索引
        if len(f.entries) > 0:
            latest_article = f.entries[0]
            if latest_article.title != last_item or last_item == "":
                with open("Tagesschau.txt", "a", encoding="utf-8") as myfile:
                    myfile.write(f"""
Article nr.: {news_nr} - Saved at: Date {today_formatted}, Time {time_formatted}\n
Published: {latest_article.published}
Title: {latest_article.title}
Text: {latest_article.description}
Link: {latest_article.link}\n
                    """)
                news_nr += 1
                articles_found += 1
                last_item = latest_article.title
        else:
            print(f"Time: [{time_formatted}] - Iteration nr. {x} - Warning: Empty RSS response, skip this check")
    except Exception as e:
        print(f"Time: [{time_formatted}] - Iteration nr. {x} - Request failed: {str(e)}, retry in next round")

    print(f"Time: [{time_formatted}] - Iteration nr. {x} - Articles found: {articles_found}")
    x += 1
    time.sleep(15)
修复说明
  • 新增全流程异常捕获,单次拉取、解析失败不会终止程序,仅打印错误信息,等待15秒后进入下一轮检测
  • 访问文章条目前增加列表长度校验,从根源上避免索引越界错误
  • 文件写入指定utf-8编码,兼容德语言特殊字符,避免编码报错
  • 移除冗余的手动关文件代码,将RSS地址提取为常量,后续维护更方便
  • 缓存最新文章对象,避免重复访问f.entries[0],减少冗余逻辑

内容的提问来源于stack exchange,提问作者Thore Wagner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 13:24:22