You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MongoDB双索引去重插入:匹配attr_name时追加values数组

MongoDB批量插入时追加不重复数组元素的实现

场景说明

我有一个MongoDB集合,数据通过insert_many插入,结构如下:

[
    {"attr_name": "a", "values": [
        {"value": "1", "embedding": [1,2,3]},
        {"value": "2", "embedding": [2,3,4]},
        {"value": "3", "embedding": [3,4,5]},
    ]},
    {"attr_name": "b", "values": [
        {"value": "1", "embedding": [1,2,3]},
        {"value": "4", "embedding": [4,5,6]},
    ]},
    {"attr_name": "c", "values": [
        {"value": "6", "embedding": [6,7,8]},
        {"value": "7", "embedding": [7,8,9]},
    ]},
]

已创建联合唯一索引避免attr_name和value重复:

collection.create_index(["attr_name", "value"], unique=True)

当前用insert_many插入新数据时,只要存在匹配的attr_name就会跳过整条记录,但我需要的是:当匹配到attr_name时,将新数据中不重复的value项追加到对应文档的values数组中;如果attr_name不存在,则直接插入整条记录。

示例

  • 现有数据:
[
    {"attr_name": "a", "values": [
        {"value": "1", "embedding": [1,2,3]},
        {"value": "2", "embedding": [2,3,4]},
    ]},
    {"attr_name": "b", "values": [
        {"value": "1", "embedding": [1,2,3]},
        {"value": "4", "embedding": [4,5,6]},
    ]},
]
  • 待插入新数据:
[
    {"attr_name": "a", "values": [
        {"value": "1", "embedding": [1,2,3]},
        {"value": "5", "embedding": [5,6,7]},
        {"value": "6", "embedding": [6,7,8]},
    ]},
    {"attr_name": "c", "values": [
        {"value": "6", "embedding": [6,7,8]},
        {"value": "7", "embedding": [7,8,9]},
    ]},
]
  • 期望最终数据状态:
[
    {"attr_name": "a", "values": [
        {"value": "1", "embedding": [1,2,3]},
        {"value": "2", "embedding": [2,3,4]},
        {"value": "5", "embedding": [5,6,7]},
        {"value": "6", "embedding": [6,7,8]},
    ]},
    {"attr_name": "b", "values": [
        {"value": "1", "embedding": [1,2,3]},
        {"value": "4", "embedding": [4,5,6]},
    ]},
    {"attr_name": "c", "values": [
        {"value": "6", "embedding": [6,7,8]},
        {"value": "7", "embedding": [7,8,9]},
    ]},
]

解决方案

不能直接用insert_many,需要改用bulk_write结合UpdateOne操作,利用$addToSet和$each实现数组去重追加,同时用upsert: True处理新attr_name的插入。

代码实现

from pymongo import MongoClient, UpdateOne

client = MongoClient("mongodb://localhost:27017/")
db = client["your_database"]
collection = db["your_collection"]

new_data = [
    {"attr_name": "a", "values": [
        {"value": "1", "embedding": [1,2,3]},
        {"value": "5", "embedding": [5,6,7]},
        {"value": "6", "embedding": [6,7,8]},
    ]},
    {"attr_name": "c", "values": [
        {"value": "6", "embedding": [6,7,8]},
        {"value": "7", "embedding": [7,8,9]},
    ]},
]

operations = []
for item in new_data:
    operations.append(
        UpdateOne(
            {"attr_name": item["attr_name"]},
            {"$addToSet": {"values": {"$each": item["values"]}}},
            upsert=True
        )
    )

# 执行批量更新
result = collection.bulk_write(operations)
print(f"插入/更新文档数: {result.upserted_count + result.modified_count}")

说明

  1. $addToSet + $each:$addToSet确保数组不会添加重复元素,配合$each可一次性追加多个元素,重复判断基于整个子文档(value和embedding完全匹配才视为重复)。
  2. upsert: True:当匹配的attr_name不存在时,自动插入包含该attr_name和values数组的新文档。
  3. 联合索引作用:原有的["attr_name", "value"]联合唯一索引从数据库层面保证attr_name和value的组合唯一,防止其他操作插入重复数据。

补充:仅基于value去重的情况

如果只需以value作为去重依据(忽略embedding变化),可逐个检查value是否存在后再追加:

operations = []
for item in new_data:
    for value_item in item["values"]:
        operations.append(
            UpdateOne(
                {"attr_name": item["attr_name"], "values.value": {"$ne": value_item["value"]}},
                {"$push": {"values": value_item}},
                upsert=True
            )
        )

collection.bulk_write(operations)

内容的提问来源于stack exchange,提问作者theodre7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 23:46:03