如何使用Nest库向ElasticSearch批量插入文档且已存在时不更新
Hey there! When working with Elasticsearch and NEST, achieving bulk inserts that only add new documents (no updates when they already exist) is totally doable—you just need to use the right operation type and configure your bulk request properly. Let me walk you through the two main approaches, depending on how much data you're dealing with.
1. For Large Datasets: Use BulkAll
BulkAll is NEST's go-to for handling large volumes of data efficiently, as it automatically batches your documents and handles retries. The key here is to use the Create operation instead of Index (since Index will overwrite existing docs).
Here's a concrete example:
var documents = GetYourDocumentCollection(); // Replace with your actual data var bulkAllObservable = client.BulkAll(documents, b => b .Index("your-index-name") .Type("_doc") // Or your document type, if using older Elasticsearch versions .BatchSize(1000) // Adjust based on your data size and cluster capacity .BackOffTime(TimeSpan.FromSeconds(10)) .BackOffRetries(3) .RefreshOnCompleted() .MaxDegreeOfParallelism(Environment.ProcessorCount) // This tells Elasticsearch to only create documents that don't exist yet .OperationName(OpType.Create) ) .Wait(TimeSpan.FromMinutes(10), next => { // Optional: Log progress or handle each batch's response if (!next.IsValid) { Console.WriteLine($"Batch failed: {next.ServerError.Error.Reason}"); } });
What's happening here?
OpType.Create: This ensures Elasticsearch will only insert the document if no document with the same ID exists. If a document with the ID is already present, Elasticsearch will return aversion_conflict_engine_exceptionfor that specific document, but the rest of the batch will still process.- If you want to suppress those conflict errors (since they're expected), you can add
.RequestConfiguration(r => r.IgnoreErrors(409))to theBulkAllconfiguration—this tells NEST to ignore 409 Conflict responses, so they won't mark the entire batch as failed.
2. For Smaller Batches: Use the Bulk API Directly
If you're dealing with smaller datasets and want more control over individual operations, you can build a bulk request manually with Create operations:
var bulkRequest = new BulkRequest("your-index-name") { Operations = documents.Select(doc => new BulkCreateOperation<YourDocumentType>(doc)).ToList() }; var bulkResponse = client.Bulk(bulkRequest); // Optional: Check which documents succeeded/failed if (!bulkResponse.IsValid) { foreach (var item in bulkResponse.ItemsWithErrors) { Console.WriteLine($"Document ID {item.Id} failed: {item.Error.Reason}"); // If the error is a conflict, it means the document already existed if (item.Error.Type == "version_conflict_engine_exception") { Console.WriteLine($"Document {item.Id} already exists—skipped."); } } }
Key Notes
- Document IDs Matter: Elasticsearch uses the document's ID to check existence. Make sure your documents have consistent, unique IDs that you can rely on to detect duplicates.
- Conflict Handling: If you don't care about logging skipped documents, adding
.IgnoreErrors(409)to your bulk request will clean up the response by ignoring those expected conflicts. - Performance: For very large datasets,
BulkAllis preferable because it handles batching and retries out of the box, reducing the risk of overwhelming your cluster.
That's it! Using Create operations instead of Index is the core trick here, and NEST gives you all the tools to handle bulk inserts while skipping updates for existing docs.
内容的提问来源于stack exchange,提问作者Fereshteh Rabet

