Elasticsearch升级至6.x:添加Join类型无需更新十亿级文档方案咨询
Great question—handling a billion documents without touching each one directly is a huge pain point when migrating to Elasticsearch 6.x's join field for parent-child relationships. Let's break down the most practical, low-impact solutions:
Option 1: Zero-Downtime Reindexing with Index Aliases
This is the gold standard for large-scale schema changes because it avoids the overhead of updating every document individually (which would be a delete+insert operation for each entry, crippling your cluster). Here's how to do it:
Create a new target index with the join field mapping
Copy your existing parent index's mapping, then add thejoinfield definition. For example:PUT /new_parent_index { "mappings": { "_doc": { "properties": { "my_join_field": { "type": "join", "relations": { "parent": ["child"] } }, // Include all existing field mappings from your old parent index here "user_id": { "type": "keyword" }, "content": { "type": "text" } } } }, "settings": { "refresh_interval": "-1" // Disable refresh during reindex for speed } }Reindex data from the old index to the new one
Use the_reindexAPI to copy documents, and inject themy_join_fieldvalue for parent documents automatically via a script:POST _reindex { "source": { "index": "old_parent_index" }, "dest": { "index": "new_parent_index" }, "script": { "source": "ctx._source.my_join_field = { 'name': 'parent' }" } }For very large datasets, add a
scrollparameter or split the reindex into batches using theslicefeature to avoid overwhelming your cluster.Switch traffic to the new index using aliases
If your application uses an alias (e.g.,parent_index) to access the original index, swap the alias to point to the new index once reindexing is complete:POST /_aliases { "actions": [ { "remove": { "index": "old_parent_index", "alias": "parent_index" } }, { "add": { "index": "new_parent_index", "alias": "parent_index" } } ] }Your application won't notice any downtime, and you can safely delete the old index once you've verified everything works.
Option 2: Alternative Parent-Child Patterns (If Reindexing Isn't Feasible)
If reindexing is temporarily off the table, consider these workarounds (though they're less ideal than using the native join field):
- Keep using separate parent/child indexes (like in 5.x): 6.x still supports this pattern via the
has_parent/has_childqueries, though it's less efficient than the join field. - Denormalize data: If your query patterns allow it, embed child documents directly into parent documents as nested fields. This avoids join overhead entirely, but only works if your data doesn't require frequent updates to child entries.
Key Notes
- Avoid updating individual documents: For 1B entries, this would generate massive amounts of cluster load and take far longer than reindexing.
- Optimize reindex performance: Disable refresh on the target index, use bulk operations, and consider adding replicas after reindexing completes to speed up the process.
内容的提问来源于stack exchange,提问作者Saurabh Johari

