CosmoDb BulkImporter抛出InvalidPartitionException异常原因咨询
This error typically pops up when your BulkExecutor instance is working with stale partition range metadata, especially during high-throughput imports that trigger collection scaling. Let's break down the causes and fixes:
Common Causes
- Stale Partition Range Metadata: When you initialize the BulkExecutor, it fetches the current partition range information for your Cosmos DB collection. If your collection scales out (splits into new partitions) while the import is running (which is highly likely at 5000 docs/sec), the existing BulkExecutor doesn't know about the new partition ranges. This mismatch leads to the "Partition range id 0 does not exist" error because the old metadata no longer reflects the current state of the collection.
- Outdated SDK Version: You're using version 1.22.0 of the DocumentDB SDK, which is quite old. Newer versions of the Cosmos DB SDK (both v2 and v3) include improvements to handle partition splits automatically, reducing the likelihood of this error occurring in the first place.
Fixes to Try
1. Re-initialize BulkExecutor on Exception (As Recommended in the Error Message)
The error explicitly suggests retrying after re-initializing the BulkExecutor instance. You can implement a retry loop that recreates the BulkExecutor when this exception is thrown. Here's an example of how to adjust your code:
int retryCount = 3; while (retryCount > 0) { try { // Re-initialize BulkExecutor if it's null or stale if (_bulkExecutor == null) { _bulkExecutor = new BulkExecutor(documentClient, collectionUri, collectionResourceId); await _bulkExecutor.InitializeAsync(); } var response = await _bulkExecutor.BulkImportAsync(data, true); break; // Exit loop once import succeeds } catch (Microsoft.Azure.Documents.InvalidPartitionException) { retryCount--; // Dispose the old instance to clear stale metadata _bulkExecutor?.Dispose(); _bulkExecutor = null; if (retryCount == 0) throw; // Re-throw if retries are exhausted } }
2. Upgrade to a Newer Cosmos DB SDK Version
Older SDK versions like 1.22.0 lack the automatic metadata refresh capabilities present in newer releases. Upgrading to the latest stable version (either v2 or v3) will help the BulkExecutor handle partition splits more gracefully without manual reinitialization in most cases.
3. Adjust Throughput (If Needed)
While 5000 docs/sec is manageable, if your collection is scaling very frequently during imports, temporarily increasing the provisioned throughput can reduce the number of splits. This stabilizes the partition ranges and minimizes the need for repeated BulkExecutor reinitialization.
Content of the question originated from Stack Exchange, asked by Toan Nguyen

