如何向MongoDB Atlas添加RSS数据?解决去重与KeyError问题
解决RSS爬虫写入MongoDB的两个问题
针对你遇到的这两个问题,我来一步步帮你搞定:
1. 最优方式写入MongoDB并避免重复数据
要避免重复写入,最可靠高效的方案是基于唯一字段创建索引,再配合MongoDB的upsert(不存在则插入,存在则更新)机制。
新闻的Link字段天然是唯一标识(同一篇新闻不会有不同链接),我们先给它创建唯一索引:
# 确保Link字段唯一,重复数据会被自动拦截 mycol.create_index("Link", unique=True)
然后替换原来的insert_many逻辑,改用批量更新/插入操作——这种方式比先查询再判断的效率高得多,适合你的定时任务场景:
# 准备批量操作队列 operations = [] for item in data: operations.append( pymongo.UpdateOne( {"Link": item["Link"]}, # 用Link作为唯一判断条件 {"$set": item}, # 存在则更新字段,不存在则插入整个文档 upsert=True ) ) # 执行批量操作 if operations: result = mycol.bulk_write(operations) print(f"处理完成:插入新文档{result.upserted_count}条,更新已有文档{result.modified_count}条")
2. 解决KeyError: 'description'或KeyError: 'pubDate'问题
不同RSS源的字段规范不统一,有些源可能缺失description或pubDate,直接用article["key"]就会触发KeyError。解决方法是用字典的get方法,它允许你指定键不存在时的默认值:
# 安全提取字段,缺失时返回默认值 title = article.get("title", "无标题") link = article.get("link", "") description = article.get("description", "无描述") pubdate = article.get("pubDate", dt.strftime("%Y-%m-%d %H:%M:%S")) # 用当前UTC时间补全缺失的发布时间
修改后的完整代码
把上面的优化整合到你的代码里,最终版本如下:
import feedparser import datetime import pymongo import json """ Crawler for the RSS. Extract the information from different RSS feeds and adds them to a MongoDB server """ # RSS源列表 feeds = [ "http://rss.nytimes.com/services/xml/rss/nyt/HomePage.xml", 'http://feeds.foxnews.com/foxnews/most-popular', 'http://www.wsj.com/xml/rss/3_7041.xml', 'http://www.wsj.com/xml/rss/3_7014.xml', 'http://www.wsj.com/xml/rss/3_7085.xml', 'http://feeds.washingtonpost.com/rss/national', 'http://rss.cnn.com/rss/cnn_topstories.rss', 'http://rss.cnn.com/rss/cnn_us.rss', 'http://feeds.feedburner.com/breitbart', 'http://www.cnbc.com/id/100003114/device/rss/rss.html', 'http://feeds.abcnews.com/abcnews/topstories', 'http://feeds.bbci.co.uk/news/rss.xml', 'https://www.wired.com/feed/', 'http://rss.upi.com/news/top_news.rss', 'http://feeds.reuters.com/reuters/topNews', 'http://rssfeeds.usatoday.com/usatoday-NewsTopStories', ] data = [] dt = datetime.datetime.utcnow() for source_url in feeds: feed = feedparser.parse(source_url) # 增加RSS解析失败的容错 if not feed.get('feed'): print(f"解析失败,跳过RSS源: {source_url}") continue source_title = feed['feed']['title'] for entry in feed["entries"]: article = json.loads(json.dumps(entry, default=str)) # 安全提取所有需要的字段 title = article.get("title", "无标题") link = article.get("link", "") description = article.get("description", "无描述") pubdate = article.get("pubDate", dt.strftime("%Y-%m-%d %H:%M:%S")) d = { "Datetime": dt, "Title": title, "Link": link, "Source": source_title, "Description": description, "PubDate": pubdate } data.append(d) # 连接MongoDB client = pymongo.MongoClient("mongodb_server") mydb = client["master_database"] mycol = mydb["news"] # 创建唯一索引 mycol.create_index("Link", unique=True) # 批量写入/更新 operations = [] for item in data: operations.append( pymongo.UpdateOne( {"Link": item["Link"]}, {"$set": item}, upsert=True ) ) if operations: result = mycol.bulk_write(operations) print(f"完成!插入新文档: {result.upserted_count}, 更新已有文档: {result.modified_count}") else: print("没有新数据需要处理")
额外优化说明:
- 把循环变量
source改成source_url,避免和源标题变量重名导致覆盖 - 增加了RSS源解析失败的判断,避免无效源导致程序崩溃
- 给所有可能缺失的字段都设置了合理默认值,保证数据完整性
内容的提问来源于stack exchange,提问作者Felisep
相关产品推荐
相关产品推荐

