MongoDB双索引去重插入:匹配attr_name时追加values数组
MongoDB批量插入时追加不重复数组元素的实现
场景说明
我有一个MongoDB集合,数据通过insert_many插入,结构如下:
[ {"attr_name": "a", "values": [ {"value": "1", "embedding": [1,2,3]}, {"value": "2", "embedding": [2,3,4]}, {"value": "3", "embedding": [3,4,5]}, ]}, {"attr_name": "b", "values": [ {"value": "1", "embedding": [1,2,3]}, {"value": "4", "embedding": [4,5,6]}, ]}, {"attr_name": "c", "values": [ {"value": "6", "embedding": [6,7,8]}, {"value": "7", "embedding": [7,8,9]}, ]}, ]
已创建联合唯一索引避免attr_name和value重复:
collection.create_index(["attr_name", "value"], unique=True)
当前用insert_many插入新数据时,只要存在匹配的attr_name就会跳过整条记录,但我需要的是:当匹配到attr_name时,将新数据中不重复的value项追加到对应文档的values数组中;如果attr_name不存在,则直接插入整条记录。
示例
- 现有数据:
[ {"attr_name": "a", "values": [ {"value": "1", "embedding": [1,2,3]}, {"value": "2", "embedding": [2,3,4]}, ]}, {"attr_name": "b", "values": [ {"value": "1", "embedding": [1,2,3]}, {"value": "4", "embedding": [4,5,6]}, ]}, ]
- 待插入新数据:
[ {"attr_name": "a", "values": [ {"value": "1", "embedding": [1,2,3]}, {"value": "5", "embedding": [5,6,7]}, {"value": "6", "embedding": [6,7,8]}, ]}, {"attr_name": "c", "values": [ {"value": "6", "embedding": [6,7,8]}, {"value": "7", "embedding": [7,8,9]}, ]}, ]
- 期望最终数据状态:
[ {"attr_name": "a", "values": [ {"value": "1", "embedding": [1,2,3]}, {"value": "2", "embedding": [2,3,4]}, {"value": "5", "embedding": [5,6,7]}, {"value": "6", "embedding": [6,7,8]}, ]}, {"attr_name": "b", "values": [ {"value": "1", "embedding": [1,2,3]}, {"value": "4", "embedding": [4,5,6]}, ]}, {"attr_name": "c", "values": [ {"value": "6", "embedding": [6,7,8]}, {"value": "7", "embedding": [7,8,9]}, ]}, ]
解决方案
不能直接用insert_many,需要改用bulk_write结合UpdateOne操作,利用$addToSet和$each实现数组去重追加,同时用upsert: True处理新attr_name的插入。
代码实现
from pymongo import MongoClient, UpdateOne client = MongoClient("mongodb://localhost:27017/") db = client["your_database"] collection = db["your_collection"] new_data = [ {"attr_name": "a", "values": [ {"value": "1", "embedding": [1,2,3]}, {"value": "5", "embedding": [5,6,7]}, {"value": "6", "embedding": [6,7,8]}, ]}, {"attr_name": "c", "values": [ {"value": "6", "embedding": [6,7,8]}, {"value": "7", "embedding": [7,8,9]}, ]}, ] operations = [] for item in new_data: operations.append( UpdateOne( {"attr_name": item["attr_name"]}, {"$addToSet": {"values": {"$each": item["values"]}}}, upsert=True ) ) # 执行批量更新 result = collection.bulk_write(operations) print(f"插入/更新文档数: {result.upserted_count + result.modified_count}")
说明
$addToSet+$each:$addToSet确保数组不会添加重复元素,配合$each可一次性追加多个元素,重复判断基于整个子文档(value和embedding完全匹配才视为重复)。upsert: True:当匹配的attr_name不存在时,自动插入包含该attr_name和values数组的新文档。- 联合索引作用:原有的
["attr_name", "value"]联合唯一索引从数据库层面保证attr_name和value的组合唯一,防止其他操作插入重复数据。
补充:仅基于value去重的情况
如果只需以value作为去重依据(忽略embedding变化),可逐个检查value是否存在后再追加:
operations = [] for item in new_data: for value_item in item["values"]: operations.append( UpdateOne( {"attr_name": item["attr_name"], "values.value": {"$ne": value_item["value"]}}, {"$push": {"values": value_item}}, upsert=True ) ) collection.bulk_write(operations)
内容的提问来源于stack exchange,提问作者theodre7
相关产品推荐
相关产品推荐

