如何删除ElasticSearch中字段数≥1000的所有文档?
Hey there! I get what you're trying to do—clean up those docs with too many fields so you can sync without hitting the index.mapping.total_fields.limit error. Your initial idea to use _delete_by_query is on the right track, but Elasticsearch doesn't have a built-in fields field that tracks the count of fields per document. We need to adjust the query to calculate that count dynamically with a script.
Here's the Correct Approach
Elasticsearch lets you use script queries to evaluate custom logic against each document. We can use a script to count the number of fields in the document's source (_source) and filter for docs where that count is ≥1000.
Step 1: Configure the _delete_by_query Request in Postman
- Request Method: POST
- URL:
http://your-es-host:9200/your-target-index/_delete_by_query(replace with your dev ES host and index name; use commas to target multiple indexes, e.g.,index1,index2) - Request Body (raw JSON):
{ "query": { "script": { "script": "doc['_source'].size() >= 1000" } }, "refresh": true // Optional: Ensures changes are visible immediately after deletion }
Key Details to Know
- How the Script Works:
doc['_source'].size()returns the number of top-level fields in the document's source data. This counts exactly the fields that contribute to thetotal_fieldslimit, so it's perfect for your use case. - Performance Considerations: Script queries run against every document in the index, so if you're dealing with a large dataset, this might take some time. For big indexes, consider adding
wait_for_completion=falseto run the operation asynchronously (you'll get a task ID to check progress later). - Verify the Deletion: After running the delete, confirm no docs with ≥1000 fields remain using a search query:
If the{ "query": { "script": { "script": "doc['_source'].size() >= 1000" } } }hits.total.valueis 0, you've successfully removed all problematic docs.
Optional: Fine-Tune the Query
If you want to exclude specific fields (though system fields don't count towards total_fields), you can adjust the script to filter out certain keys. For example, to exclude a field named internal_metadata:
{ "query": { "script": { "script": "doc['_source'].keySet().stream().filter(key -> !key.equals('internal_metadata')).count() >= 1000" } } }
内容的提问来源于stack exchange,提问作者cmac

