在Express中基于documentId删除Azure Cognitive Search索引文档
问题描述
我在Azure Cognitive Search中有一个名为document的索引,文档结构示例如下:
{"@search.score": 1,"id": "3412b974-ca9e-4bce-a592-78084a71e4c6","content": null,"documentId": "3d8c0912-e292-483f-80a2-5e63c57b9504","bucket": "test-storage-autonomize","name": "demo-project/75131_Patient.pdf","metadata": "","form": [],"table": [],"medicalEntities": [{"BeginOffset": 1961,"Category": "MEDICAL_CONDITION","EndOffset": 1964,"Id": 56,"Score": 0.808318018913269,"Text": "ARF","Type": "DX_NAME","Attributes": [],"ICD10CMConcepts": [{"Code": "J96.0","Description": "Acute respiratory failure","Score": 0.8527929592132568"},{"Code": "J96.0","Description": "Acute respiratory failure","Score": 0.8527929592132568"},{"Code": "J96.2","Description": "Acute and chronic respiratory failure","Score": 0.8522305774688721"},{"Code": "N17.2","Description": "Acute kidney failure with medullary necrosis","Score": 0.8506965065002441"},{"Code": "N17.2","Description": "Acute renal failure with medullary necrosis","Score": 0.8506684684753418"}],"RxNormConcepts": [],"CPT_Current_Procedural_Terminology": [],"Traits": [{"Name": "SIGN","Score": 0.808318018913269}]}]}
需要基于documentId删除索引中的文档,但documentId不是索引主键(主键是id)。目前在Express框架中使用Azure SDK尝试过两种方式:
- 先通过
documentId搜索获取文档再删除,存在额外开销; - 尝试
action.delete方法,但该方法需要传入主键,无法直接用documentId。
可行解决方案
1. 使用批量删除的过滤条件(推荐)
Azure Cognitive Search的批量删除支持通过filter参数匹配非主键字段,一次性删除符合条件的文档,无需先查询再删除。
前提是documentId字段已设置为可过滤(filterable),如果未配置,需要先更新索引的字段定义,将documentId的filterable属性设为true。
Express中使用Azure SDK的示例代码:
const { SearchIndexClient, AzureKeyCredential, IndexDocumentsClient } = require("@azure/search-documents"); async function deleteByDocumentId(serviceEndpoint, apiKey, indexName, targetDocumentId) { const client = new SearchIndexClient(serviceEndpoint, new AzureKeyCredential(apiKey)); const indexDocumentsClient = client.getIndexDocumentsClient(indexName); // 构建带过滤条件的删除请求 const deleteOptions = { filter: `documentId eq '${targetDocumentId}'` }; const result = await indexDocumentsClient.deleteDocuments(deleteOptions); return result; } // 集成到Express路由 app.delete('/documents/:documentId', async (req, res) => { try { await deleteByDocumentId( process.env.AZURE_SEARCH_ENDPOINT, process.env.AZURE_SEARCH_API_KEY, 'document', req.params.documentId ); res.status(200).json({ message: '文档删除成功' }); } catch (error) { res.status(500).json({ error: error.message }); } });
2. 调整索引主键(如果业务允许)
如果业务场景支持,可以将索引主键从id修改为documentId,后续即可直接通过action.delete方法传入documentId完成删除。但此操作需要重建索引,且需确保documentId全局唯一,适合索引创建初期或可接受数据迁移的场景。
3. 优化查询后删除的方式(备选)
如果暂时无法调整索引配置或使用批量过滤删除,可以优化查询步骤,仅返回主键id字段,减少数据传输开销:
async function getPrimaryKeyByDocumentId(serviceEndpoint, apiKey, indexName, targetDocumentId) { const client = new SearchIndexClient(serviceEndpoint, new AzureKeyCredential(apiKey)); const searchClient = client.getSearchClient(indexName); const searchResult = await searchClient.search(`documentId eq '${targetDocumentId}'`, { select: 'id', // 仅返回主键字段 top: 1 // 假设documentId唯一,只取第一个结果 }); let primaryKey = null; for await (const result of searchResult.results) { primaryKey = result.document.id; break; } return primaryKey; } async function deleteDocumentById(serviceEndpoint, apiKey, indexName, primaryKey) { const client = new SearchIndexClient(serviceEndpoint, new AzureKeyCredential(apiKey)); const indexDocumentsClient = client.getIndexDocumentsClient(indexName); return await indexDocumentsClient.deleteDocuments([{ id: primaryKey }]); } // 路由中调用 app.delete('/documents/:documentId', async (req, res) => { try { const primaryKey = await getPrimaryKeyByDocumentId( process.env.AZURE_SEARCH_ENDPOINT, process.env.AZURE_SEARCH_API_KEY, 'document', req.params.documentId ); if (!primaryKey) { return res.status(404).json({ message: '未找到对应文档' }); } await deleteDocumentById( process.env.AZURE_SEARCH_ENDPOINT, process.env.AZURE_SEARCH_API_KEY, 'document', primaryKey ); res.status(200).json({ message: '文档删除成功' }); } catch (error) { res.status(500).json({ error: error.message }); } });
内容的提问来源于stack exchange,提问作者ctaneja
相关产品推荐
相关产品推荐

