如何用Python 2.7(feedparser库)持续监控RSS/Atom Feed并同步至数据库?
Got it, let's walk through building this feed monitoring system step by step. I'll stick to Python 2.7, feedparser, and use SQLite for the database (it's lightweight, built-in, and perfect for this use case—you can swap it for MySQL/PostgreSQL later if needed).
Step 1: Set Up Dependencies
First, install feedparser (since Python 2.7 typically uses pip2):
pip2 install feedparser
We'll use Python's built-in sqlite3 for the database, so no extra installs are needed there.
Step 2: Database Setup
We need a table to track which articles we've already seen. The key is storing a unique identifier for each article—most feeds use an id field, but if that's missing, the article's link works as a fallback.
Here's the table schema we'll use:
CREATE TABLE IF NOT EXISTS feed_entries ( id INTEGER PRIMARY KEY AUTOINCREMENT, feed_url TEXT NOT NULL, entry_identifier TEXT UNIQUE NOT NULL, -- This is our unique check title TEXT NOT NULL, link TEXT NOT NULL, published_date TEXT, summary TEXT );
Step 3: Core Logic Implementation
Let's put it all together with code that:
- Connects to the database
- Parses the feed
- Checks if each article is already in the DB
- Adds new articles and triggers an alert
- Runs on a loop to keep monitoring
import feedparser import sqlite3 import time # Configuration FEED_URL = "https://example.com/feed" # Replace with your target feed URL CHECK_INTERVAL = 1800 # Check every 30 minutes (in seconds) DB_PATH = "feed_monitor.db" def init_db(): """Initialize the database and create table if it doesn't exist""" conn = sqlite3.connect(DB_PATH) cursor = conn.cursor() cursor.execute(''' CREATE TABLE IF NOT EXISTS feed_entries ( id INTEGER PRIMARY KEY AUTOINCREMENT, feed_url TEXT NOT NULL, entry_identifier TEXT UNIQUE NOT NULL, title TEXT NOT NULL, link TEXT NOT NULL, published_date TEXT, summary TEXT ) ''') conn.commit() conn.close() def entry_exists(entry_id): """Check if an article is already in the database""" conn = sqlite3.connect(DB_PATH) cursor = conn.cursor() cursor.execute("SELECT 1 FROM feed_entries WHERE entry_identifier = ?", (entry_id,)) exists = cursor.fetchone() is not None conn.close() return exists def add_entry(feed_url, entry): """Add a new article to the database and trigger alert""" # Use entry.id as unique ID, fall back to link if id is missing entry_id = entry.get('id', entry.link) published = entry.get('published', 'Unknown') summary = entry.get('summary', '') conn = sqlite3.connect(DB_PATH) cursor = conn.cursor() try: cursor.execute(''' INSERT INTO feed_entries (feed_url, entry_identifier, title, link, published_date, summary) VALUES (?, ?, ?, ?, ?, ?) ''', (feed_url, entry_id, entry.title, entry.link, published, summary)) conn.commit() print("✅ Added new article: {}".format(entry.title)) # Trigger your alert here—this is a basic print example print("⚠️ ALERT: New article published!") print("Title: {}".format(entry.title)) print("Link: {}\n".format(entry.link)) except sqlite3.IntegrityError: # Safety check in case entry_exists missed something print("⚠️ Article already exists in DB: {}\n".format(entry.title)) finally: conn.close() def monitor_feed(): """Main loop to continuously monitor the feed""" init_db() print("Starting feed monitor for: {}".format(FEED_URL)) print("Checking every {} seconds...\n".format(CHECK_INTERVAL)) while True: try: # Parse the feed feed = feedparser.parse(FEED_URL) if feed.bozo != 0: print("❌ Error parsing feed: {}\n".format(feed.bozo_exception)) time.sleep(CHECK_INTERVAL) continue # Process entries (reverse to handle oldest first, optional) for entry in reversed(feed.entries): entry_id = entry.get('id', entry.link) if not entry_exists(entry_id): add_entry(FEED_URL, entry) print("✅ Check completed. Waiting for next interval...\n") time.sleep(CHECK_INTERVAL) except Exception as e: print("❌ Unexpected error: {}\n".format(str(e))) time.sleep(CHECK_INTERVAL) if __name__ == "__main__": monitor_feed()
Key Notes & Customizations
- Unique Identifier: If your target feed doesn't have a reliable
idfield, swapentry.idwithentry.link(most article links are unique). - Alerting: The example uses
printfor alerts, but you can expand this to:- Send emails with Python's built-in
smtplib - Post to Slack/Teams via their APIs
- Trigger desktop notifications (install
plyerwithpip2 install plyerfor this)
- Send emails with Python's built-in
- Multiple Feeds: To monitor multiple feeds, create a list of feed URLs and loop through them in the
monitor_feedfunction. - Error Handling: The code includes basic error handling, but you can add specific checks (like network timeouts) around the
feedparser.parsecall. - Python 2.7 Quirks: Remember that
printis a statement (not a function) in Python 2.7, so we useprint "text"instead ofprint("text").
Testing It Out
- Replace
FEED_URLwith your target blog's RSS/Atom feed. - Run the script:
python2 feed_monitor.py - Publish a new article on the blog—you’ll see the alert print out, and the article will be stored in the SQLite database.
内容的提问来源于stack exchange,提问作者Syed Aashir

