如何向MongoDB插入数据时避免产生重复记录
MongoDB插入避免重复记录解决方案
你原代码存在三个核心问题:
update()方法参数传递错误,缺少查询条件,也没有指定更新模式- 把MongoDB连接创建放在for循环内部,会产生大量无效连接,严重影响性能
- 字段拼写错误,把
country写为了contry
具体实现方案
1. 提前创建唯一索引
首先给集合的业务唯一键(这里你可以用id字段作为唯一判断标准)创建唯一索引,从数据库层面兜底阻止重复数据插入:
# 只需要执行一次,后续不用重复创建 collection.create_index("id", unique=True)
如果你的业务判断重复的逻辑是多个字段组合唯一,也可以创建联合唯一索引:
collection.create_index([("id", 1), ("source_url", 1)], unique=True)
2. 使用upsert模式的更新操作
MongoDB的update_one方法支持upsert参数,设置为True时会自动判断:匹配到查询条件的记录就更新内容,没有匹配到就插入新记录,刚好满足你修改文件后覆盖旧数据、不产生重复的需求。
3. 修正后的完整代码
import csv from pymongo import MongoClient # 数据库连接全局只创建一次即可 client = MongoClient('localhost', 27017) db = client['Parser'] collection = db['csv'] # 提前创建唯一索引,只需要执行一次 collection.create_index("id", unique=True) lst = [] with open('qAfpxiFfaq31nRfqWiYnKiJRW7IRiyx6.csv', encoding='utf-8') as f: reader = csv.DictReader(f) for row in reader: title_dict = { "id": row["id"], "source_name": row["source_name"], "title": row["title"], "source_url": row["source_url"], "language": row["language"], "country": row["country"], "published_at": row["published_at"] } lst.append(title_dict) # 以id为查询条件,匹配到就更新整条数据,没匹配到就插入 collection.update_one( {"id": title_dict["id"]}, {"$set": title_dict}, upsert=True )
内容的提问来源于stack exchange,提问作者makim
相关产品推荐
相关产品推荐

