You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于新闻聚合安卓项目,如何用Python定期爬取并存储多站点RSS Feed?

实现Python定期自动爬取RSS Feed并存入MySQL的方案

看起来你已经搭建了一套不错的新闻聚合系统架构,要实现定期自动爬取RSS Feed并存库的需求,咱们可以分这几个步骤来做:

1. 选择并安装RSS解析库

Python里处理RSS Feed最常用的是feedparser库,它能轻松解析各种标准的RSS/Atom格式。安装命令很简单:

pip install feedparser

2. 编写RSS爬取核心逻辑

先写一个函数来处理单个RSS Feed的爬取,提取关键信息(标题、链接、发布时间、摘要/内容等),同时处理可能的异常:

import feedparser
import datetime

def fetch_rss_feed(feed_url):
    try:
        feed = feedparser.parse(feed_url)
        if feed.bozo != 0:
            # 解析出错,返回空列表
            print(f"解析Feed失败: {feed_url},错误信息: {feed.bozo_exception}")
            return []
        
        articles = []
        for entry in feed.entries:
            # 提取字段,处理可能缺失的情况
            article = {
                'title': entry.get('title', '无标题'),
                'link': entry.get('link', ''),
                'published': entry.get('published', entry.get('updated', '')),
                'summary': entry.get('summary', entry.get('description', '')),
                'source': feed_url  # 记录来源Feed
            }
            # 格式化发布时间(如果有)
            if article['published']:
                try:
                    article['published'] = datetime.datetime.strptime(article['published'], '%a, %d %b %Y %H:%M:%S %z').strftime('%Y-%m-%d %H:%M:%S')
                except ValueError:
                    # 兼容其他时间格式,这里可以根据需要扩展
                    article['published'] = datetime.datetime.now().strftime('%Y-%m-%d %H:%M:%S')
            articles.append(article)
        return articles
    except Exception as e:
        print(f"爬取Feed {feed_url} 时出错: {str(e)}")
        return []

3. 连接MySQL并存储数据

接下来要把爬取到的文章存入你的MySQL数据库,这里用mysql-connector-python库来连接:

pip install mysql-connector-python

然后编写数据库操作函数,重点是避免重复插入(比如根据文章链接判断是否已存在):

import mysql.connector
from mysql.connector import Error

def connect_to_db():
    try:
        connection = mysql.connector.connect(
            host='localhost',  # 你的本地服务器地址
            database='your_database_name',  # 替换成你的数据库名
            user='your_username',  # 数据库用户名
            password='your_password'  # 数据库密码
        )
        if connection.is_connected():
            return connection
    except Error as e:
        print(f"数据库连接失败: {str(e)}")
        return None

def save_articles_to_db(articles):
    connection = connect_to_db()
    if not connection:
        return
    
    cursor = connection.cursor()
    # 先检查文章是否已存在,不存在则插入
    insert_query = """
    INSERT INTO news_articles (title, link, published, summary, source)
    VALUES (%s, %s, %s, %s, %s)
    ON DUPLICATE KEY UPDATE title=VALUES(title), published=VALUES(published), summary=VALUES(summary)
    """
    # 注意:需要确保你的news_articles表中link字段设置了UNIQUE约束,这样才能触发ON DUPLICATE KEY
    
    try:
        for article in articles:
            cursor.execute(insert_query, (
                article['title'],
                article['link'],
                article['published'],
                article['summary'],
                article['source']
            ))
        connection.commit()
        print(f"成功存储/更新 {cursor.rowcount} 篇文章")
    except Error as e:
        print(f"存储数据出错: {str(e)}")
        connection.rollback()
    finally:
        if connection.is_connected():
            cursor.close()
            connection.close()

4. 实现定期自动执行

有两种常用方式来实现定期爬取:

方式一:用Python的schedule库(纯Python实现)

适合不想折腾系统定时任务的场景,安装库:

pip install schedule

然后编写调度逻辑:

import schedule
import time

# 定义你的RSS Feed列表
RSS_FEEDS = [
    'https://example.com/rss',
    'https://another-site.com/atom',
    # 添加更多Feed地址
]

def run_scraper():
    print(f"开始执行爬取任务: {datetime.datetime.now()}")
    for feed_url in RSS_FEEDS:
        articles = fetch_rss_feed(feed_url)
        if articles:
            save_articles_to_db(articles)
    print("爬取任务执行完成")

# 设置定时规则,比如每2小时执行一次
schedule.every(2).hours.do(run_scraper)

# 保持脚本运行
while True:
    schedule.run_pending()
    time.sleep(60)

方式二:用系统定时任务(更稳定)

如果你的本地服务器是Linux,推荐用crontab;如果是Windows,用「任务计划程序」。

Linux crontab设置:

  1. 打开终端,输入crontab -e编辑定时任务
  2. 添加一行规则,比如每天凌晨2点执行爬取脚本:
0 2 * * * /usr/bin/python3 /path/to/your/rss_scraper.py >> /path/to/logs/rss_scraper.log 2>&1
  • /usr/bin/python3是你的Python解释器路径(可以用which python3查看)
  • /path/to/your/rss_scraper.py是你的爬取脚本路径
  • 后面的>> ...是把日志输出到文件,方便排查问题

Windows任务计划程序:

  1. 搜索并打开「任务计划程序」
  2. 创建基本任务,设置触发时间(比如每天凌晨2点)
  3. 操作选择「启动程序」,程序或脚本选择你的Python解释器(比如python.exe),添加参数为你的脚本路径(比如C:\path\to\rss_scraper.py)

一些优化建议

  • 添加日志记录:用Python的logging模块替代print,更规范地记录爬取和存储过程中的信息
  • 设置请求头:在feedparser请求时添加User-Agent,模拟浏览器请求,避免被网站拦截:
    feedparser.parse(feed_url, agent='Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36')
    
  • 批量插入:如果爬取的文章很多,把多个文章打包成批量插入语句,能大幅提高存储效率
  • 处理字段差异:不同网站的RSS字段可能有差异,比如有些用content代替summary,可以在提取时增加兼容逻辑
  • 清理过期数据:定期删除数据库中超过一定时间的旧文章,避免数据库体积过大

内容的提问来源于stack exchange,提问作者Antenna_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:27:13