You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何向MongoDB Atlas添加RSS数据?解决去重与KeyError问题

解决RSS爬虫写入MongoDB的两个问题

针对你遇到的这两个问题,我来一步步帮你搞定:

1. 最优方式写入MongoDB并避免重复数据

要避免重复写入,最可靠高效的方案是基于唯一字段创建索引,再配合MongoDB的upsert(不存在则插入,存在则更新)机制。

新闻的Link字段天然是唯一标识(同一篇新闻不会有不同链接),我们先给它创建唯一索引:

# 确保Link字段唯一,重复数据会被自动拦截
mycol.create_index("Link", unique=True)

然后替换原来的insert_many逻辑,改用批量更新/插入操作——这种方式比先查询再判断的效率高得多,适合你的定时任务场景:

# 准备批量操作队列
operations = []
for item in data:
    operations.append(
        pymongo.UpdateOne(
            {"Link": item["Link"]},  # 用Link作为唯一判断条件
            {"$set": item},          # 存在则更新字段,不存在则插入整个文档
            upsert=True
        )
    )

# 执行批量操作
if operations:
    result = mycol.bulk_write(operations)
    print(f"处理完成:插入新文档{result.upserted_count}条,更新已有文档{result.modified_count}条")

2. 解决KeyError: 'description'或KeyError: 'pubDate'问题

不同RSS源的字段规范不统一,有些源可能缺失description或pubDate,直接用article["key"]就会触发KeyError。解决方法是用字典的get方法,它允许你指定键不存在时的默认值:

# 安全提取字段,缺失时返回默认值
title = article.get("title", "无标题")
link = article.get("link", "")
description = article.get("description", "无描述")
pubdate = article.get("pubDate", dt.strftime("%Y-%m-%d %H:%M:%S"))  # 用当前UTC时间补全缺失的发布时间

修改后的完整代码

把上面的优化整合到你的代码里,最终版本如下:

import feedparser
import datetime
import pymongo
import json

""" Crawler for the RSS. Extract the information from different RSS feeds and adds them to a MongoDB server """
# RSS源列表
feeds = [
 "http://rss.nytimes.com/services/xml/rss/nyt/HomePage.xml",
 'http://feeds.foxnews.com/foxnews/most-popular',
 'http://www.wsj.com/xml/rss/3_7041.xml',
 'http://www.wsj.com/xml/rss/3_7014.xml',
 'http://www.wsj.com/xml/rss/3_7085.xml',
 'http://feeds.washingtonpost.com/rss/national',
 'http://rss.cnn.com/rss/cnn_topstories.rss',
 'http://rss.cnn.com/rss/cnn_us.rss',
 'http://feeds.feedburner.com/breitbart',
 'http://www.cnbc.com/id/100003114/device/rss/rss.html',
 'http://feeds.abcnews.com/abcnews/topstories',
 'http://feeds.bbci.co.uk/news/rss.xml',
 'https://www.wired.com/feed/',
 'http://rss.upi.com/news/top_news.rss',
 'http://feeds.reuters.com/reuters/topNews',
 'http://rssfeeds.usatoday.com/usatoday-NewsTopStories',
]

data = []
dt = datetime.datetime.utcnow()
for source_url in feeds:
    feed = feedparser.parse(source_url)
    # 增加RSS解析失败的容错
    if not feed.get('feed'):
        print(f"解析失败,跳过RSS源: {source_url}")
        continue
    source_title = feed['feed']['title']
    for entry in feed["entries"]:
        article = json.loads(json.dumps(entry, default=str))
        # 安全提取所有需要的字段
        title = article.get("title", "无标题")
        link = article.get("link", "")
        description = article.get("description", "无描述")
        pubdate = article.get("pubDate", dt.strftime("%Y-%m-%d %H:%M:%S"))
        d = {
            "Datetime": dt,
            "Title": title,
            "Link": link,
            "Source": source_title,
            "Description": description,
            "PubDate": pubdate
        }
        data.append(d)

# 连接MongoDB
client = pymongo.MongoClient("mongodb_server")
mydb = client["master_database"]
mycol = mydb["news"]

# 创建唯一索引
mycol.create_index("Link", unique=True)

# 批量写入/更新
operations = []
for item in data:
    operations.append(
        pymongo.UpdateOne(
            {"Link": item["Link"]},
            {"$set": item},
            upsert=True
        )
    )

if operations:
    result = mycol.bulk_write(operations)
    print(f"完成!插入新文档: {result.upserted_count}, 更新已有文档: {result.modified_count}")
else:
    print("没有新数据需要处理")

额外优化说明:

  • 把循环变量source改成source_url,避免和源标题变量重名导致覆盖
  • 增加了RSS源解析失败的判断,避免无效源导致程序崩溃
  • 给所有可能缺失的字段都设置了合理默认值,保证数据完整性

内容的提问来源于stack exchange,提问作者Felisep

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:11:37