You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python批量爬取多个RSS源并将结果分别存储到对应CSV

批量爬取多RSS源并独立存储CSV实现方案

核心思路是把单源爬取逻辑封装为可复用函数,循环读取RSS源配置列表逐个执行,不需要为每个源单独写脚本。

前置准备

先整理所有RSS源链接到名为rss_source_list.csv的文件中,和爬取脚本放在同一目录,文件格式参考:

feed_id,url
feed01,https://example1.com/rss.xml
feed02,https://example2.com/feed
feed03,https://example3.com/rss

其中feed_id为每个源的唯一标识,用于生成不重名的输出CSV文件,不要重复。

完整实现代码

import os
import requests
from bs4 import BeautifulSoup
import pandas as pd

# 初始化结果存储目录,不存在则自动创建
os.makedirs('results', exist_ok=True)

def crawl_single_feed(feed_id: str, url: str, timeout: int = 10) -> None:
    """爬取单个RSS源,结果存入对应命名的CSV文件"""
    try:
        resp = requests.get(url, timeout=timeout)
        resp.raise_for_status()
        soup = BeautifulSoup(resp.text, 'html.parser')

        entries = []
        # 同时兼容Atom格式的entry标签和RSS2.0格式的item标签
        entries.extend(soup.find_all('entry'))
        entries.extend(soup.find_all('item'))

        result = []
        for entry in entries:
            # 字段提取加容错,兼容不同源的标签差异
            title = entry.find('title')
            pubdate = entry.find('published') or entry.find('pubdate') or entry.find('pubDate')
            content = entry.find('content') or entry.find('description')
            link = entry.find('link')

            item = {
                'Title': title.text.strip() if title else '',
                'Pubdate': pubdate.text.strip() if pubdate else '',
                'Content': content.text.strip() if content else '',
                'Link': link.get('href', '').strip() if link else ''
            }
            result.append(item)

        # 存储结果
        save_path = f'results/results_{feed_id}.csv'
        pd.DataFrame(result).to_csv(save_path, index=False, encoding='utf-8-sig')
        print(f"[SUCCESS] Feed {feed_id} 爬取完成,共{len(result)}条内容,存储路径:{save_path}")
    except Exception as e:
        print(f"[FAILED] Feed {feed_id} 爬取失败,错误信息:{str(e)}")

if __name__ == '__main__':
    # 读取所有RSS源配置
    feed_list = pd.read_csv('rss_source_list.csv')
    # 循环执行所有爬取任务
    for _, row in feed_list.iterrows():
        crawl_single_feed(feed_id=str(row['feed_id']), url=row['url'].strip())

补充说明

  • 代码加了异常捕获,单个源爬取失败(比如链接失效、超时、页面结构不匹配)不会中断整个批量任务,控制台会打印对应失败信息方便排查。
  • 输出文件默认用utf-8-sig编码,直接用Excel打开不会出现中文乱码。
  • 如果RSS源数量超过50个,可以自行把串行循环改为多线程/协程版本提升爬取效率,常规几十源的场景串行执行足够稳定。
  • 如果需要自定义请求头(比如部分源会拦截默认requests请求头),在requests.get里加headers参数即可。

内容的提问来源于stack exchange,提问作者Parzival

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.08 16:15:16