You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 2.7(feedparser库)持续监控RSS/Atom Feed并同步至数据库?

RSS/Atom Feed Monitor with Python 2.7 & Feedparser

Got it, let's walk through building this feed monitoring system step by step. I'll stick to Python 2.7, feedparser, and use SQLite for the database (it's lightweight, built-in, and perfect for this use case—you can swap it for MySQL/PostgreSQL later if needed).

Step 1: Set Up Dependencies

First, install feedparser (since Python 2.7 typically uses pip2):

pip2 install feedparser

We'll use Python's built-in sqlite3 for the database, so no extra installs are needed there.

Step 2: Database Setup

We need a table to track which articles we've already seen. The key is storing a unique identifier for each article—most feeds use an id field, but if that's missing, the article's link works as a fallback.

Here's the table schema we'll use:

CREATE TABLE IF NOT EXISTS feed_entries (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    feed_url TEXT NOT NULL,
    entry_identifier TEXT UNIQUE NOT NULL, -- This is our unique check
    title TEXT NOT NULL,
    link TEXT NOT NULL,
    published_date TEXT,
    summary TEXT
);

Step 3: Core Logic Implementation

Let's put it all together with code that:

  1. Connects to the database
  2. Parses the feed
  3. Checks if each article is already in the DB
  4. Adds new articles and triggers an alert
  5. Runs on a loop to keep monitoring
import feedparser
import sqlite3
import time

# Configuration
FEED_URL = "https://example.com/feed"  # Replace with your target feed URL
CHECK_INTERVAL = 1800  # Check every 30 minutes (in seconds)
DB_PATH = "feed_monitor.db"

def init_db():
    """Initialize the database and create table if it doesn't exist"""
    conn = sqlite3.connect(DB_PATH)
    cursor = conn.cursor()
    cursor.execute('''
        CREATE TABLE IF NOT EXISTS feed_entries (
            id INTEGER PRIMARY KEY AUTOINCREMENT,
            feed_url TEXT NOT NULL,
            entry_identifier TEXT UNIQUE NOT NULL,
            title TEXT NOT NULL,
            link TEXT NOT NULL,
            published_date TEXT,
            summary TEXT
        )
    ''')
    conn.commit()
    conn.close()

def entry_exists(entry_id):
    """Check if an article is already in the database"""
    conn = sqlite3.connect(DB_PATH)
    cursor = conn.cursor()
    cursor.execute("SELECT 1 FROM feed_entries WHERE entry_identifier = ?", (entry_id,))
    exists = cursor.fetchone() is not None
    conn.close()
    return exists

def add_entry(feed_url, entry):
    """Add a new article to the database and trigger alert"""
    # Use entry.id as unique ID, fall back to link if id is missing
    entry_id = entry.get('id', entry.link)
    published = entry.get('published', 'Unknown')
    summary = entry.get('summary', '')
    
    conn = sqlite3.connect(DB_PATH)
    cursor = conn.cursor()
    try:
        cursor.execute('''
            INSERT INTO feed_entries (feed_url, entry_identifier, title, link, published_date, summary)
            VALUES (?, ?, ?, ?, ?, ?)
        ''', (feed_url, entry_id, entry.title, entry.link, published, summary))
        conn.commit()
        print("✅ Added new article: {}".format(entry.title))
        # Trigger your alert here—this is a basic print example
        print("⚠️ ALERT: New article published!")
        print("Title: {}".format(entry.title))
        print("Link: {}\n".format(entry.link))
    except sqlite3.IntegrityError:
        # Safety check in case entry_exists missed something
        print("⚠️ Article already exists in DB: {}\n".format(entry.title))
    finally:
        conn.close()

def monitor_feed():
    """Main loop to continuously monitor the feed"""
    init_db()
    print("Starting feed monitor for: {}".format(FEED_URL))
    print("Checking every {} seconds...\n".format(CHECK_INTERVAL))
    
    while True:
        try:
            # Parse the feed
            feed = feedparser.parse(FEED_URL)
            if feed.bozo != 0:
                print("❌ Error parsing feed: {}\n".format(feed.bozo_exception))
                time.sleep(CHECK_INTERVAL)
                continue
            
            # Process entries (reverse to handle oldest first, optional)
            for entry in reversed(feed.entries):
                entry_id = entry.get('id', entry.link)
                if not entry_exists(entry_id):
                    add_entry(FEED_URL, entry)
            
            print("✅ Check completed. Waiting for next interval...\n")
            time.sleep(CHECK_INTERVAL)
        
        except Exception as e:
            print("❌ Unexpected error: {}\n".format(str(e)))
            time.sleep(CHECK_INTERVAL)

if __name__ == "__main__":
    monitor_feed()

Key Notes & Customizations

  • Unique Identifier: If your target feed doesn't have a reliable id field, swap entry.id with entry.link (most article links are unique).
  • Alerting: The example uses print for alerts, but you can expand this to:
    • Send emails with Python's built-in smtplib
    • Post to Slack/Teams via their APIs
    • Trigger desktop notifications (install plyer with pip2 install plyer for this)
  • Multiple Feeds: To monitor multiple feeds, create a list of feed URLs and loop through them in the monitor_feed function.
  • Error Handling: The code includes basic error handling, but you can add specific checks (like network timeouts) around the feedparser.parse call.
  • Python 2.7 Quirks: Remember that print is a statement (not a function) in Python 2.7, so we use print "text" instead of print("text").

Testing It Out

  1. Replace FEED_URL with your target blog's RSS/Atom feed.
  2. Run the script: python2 feed_monitor.py
  3. Publish a new article on the blog—you’ll see the alert print out, and the article will be stored in the SQLite database.

内容的提问来源于stack exchange,提问作者Syed Aashir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:27:01