You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用存储过程插入Azure Cosmos DB文档时分区键报错求助

问题分析

你碰到的错误Requests originating from scripts cannot reference partition keys other than the one for which client request was submitted,核心原因是存储过程是绑定到特定分区运行的:当你调用存储过程时指定了partitionKey: 'India',存储过程就只能操作该分区内的资源。而你的代码存在两个关键问题导致报错:

  • 你传入的是包含多个文档的数组,但存储过程直接把整个数组当作单个文档插入——数组本身没有Country字段,Cosmos DB无法识别它的分区键,自然会触发分区不匹配错误。
  • 存储过程没有遍历数组逐个创建文档,完全不符合批量插入的需求。
解决方案

我们需要修改存储过程代码,让它正确处理批量文档,同时确保调用逻辑符合分区限制:

1. 修复存储过程代码

更新存储过程的JavaScript逻辑,使其遍历传入的文档数组,逐个创建符合分区要求的文档:

function (docsJson) {
    var docs = JSON.parse(docsJson);
    var collection = getContext().getCollection();
    var collectionLink = collection.getSelfLink();
    var count = 0;

    // 递归遍历文档数组,逐个插入
    function createDoc(index) {
        if (index >= docs.length) {
            getContext().getResponse().setBody(count);
            return;
        }
        var doc = docs[index];
        // 校验文档是否包含必填分区键
        if (!doc.Country) {
            throw new Error('Document missing required partition key: Country');
        }
        collection.createDocument(collectionLink, doc, function(err) {
            if (err) {
                throw new Error('Failed to create document: ' + err.message);
            }
            count++;
            createDoc(index + 1);
        });
    }

    // 启动插入流程
    if (docs.length > 0) {
        createDoc(0);
    } else {
        getContext().getResponse().setBody(0);
    }
}

2. 调整Python调用逻辑

确保传入的是正确的文档数组,并且调用存储过程时指定的partitionKey和所有文档的Country字段值完全一致:

# 拆分数据为多个批次(原代码的拆分逻辑保留)
df_list = np.array_split(dataframe, 50)
start_time = time.time()
total_inserted = 0

for df_chunk in df_list:
    chunk_docs = df_chunk.to_dict('records')
    # 调用存储过程,确保partitionKey与当前批次文档的Country值一致
    insert_count = client.ExecuteStoredProcedure(
        proc_link, 
        json.dumps(chunk_docs, default=str), 
        {'partitionKey': 'India'}
    )
    total_inserted += insert_count

print(f"Successfully inserted {total_inserted} documents")
print("--- %s seconds ---" % (time.time() - start_time))

额外注意事项

  • 分区限制:存储过程只能在单个分区内运行,所以你插入的所有文档必须属于同一个分区(即Country字段值相同)。如果数据包含多个国家,需要先按Country分组,再分别调用存储过程处理每个分区的数据。
  • 吞吐量优化:你的代码中调整吞吐量的逻辑是可行的,但频繁调整可能会有延迟,建议根据实际插入数据量合理设置。
  • 错误处理:可以在存储过程和Python代码中添加更详细的错误捕获逻辑,方便排查插入失败的具体原因。
完整修改后的关键代码片段

存储过程定义部分

sproc = {
    'id': 'storedProcedure',
    'body': (
        'function (docsJson) {' +
        '    var docs = JSON.parse(docsJson);' +
        '    var collection = getContext().getCollection();' +
        '    var collectionLink = collection.getSelfLink();' +
        '    var count = 0;' +
        '' +
        '    function createDoc(index) {' +
        '        if (index >= docs.length) {' +
        '            getContext().getResponse().setBody(count);' +
        '            return;' +
        '        }' +
        '        var doc = docs[index];' +
        '        if (!doc.Country) {' +
        '            throw new Error(\'Document missing required partition key: Country\');' +
        '        }' +
        '        collection.createDocument(collectionLink, doc, function(err) {' +
        '            if (err) {' +
        '                throw new Error(\'Failed to create document: \' + err.message);' +
        '            }' +
        '            count++;' +
        '            createDoc(index + 1);' +
        '        });' +
        '    }' +
        '' +
        '    if (docs.length > 0) {' +
        '        createDoc(0);' +
        '    } else {' +
        '        getContext().getResponse().setBody(0);' +
        '    }' +
        '}'
    )
}

这样修改后,存储过程会正确遍历每个文档并插入,同时确保所有操作都在指定的分区内,就不会再触发分区不匹配的错误了。

内容的提问来源于stack exchange,提问作者Anurag

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 00:52:49