Elasticsearch源索引持续写入时重新索引的行为及生产环境使用Reindex API的相关问询
Elasticsearch Reindex API 持续写入源索引时的行为解析
Great question—this is a critical scenario for production environments, so let’s break down exactly what happens and how to handle it.
一、源索引持续写入时的Reindex行为表现
When you kick off a _reindex job while your source index is still receiving writes, here’s what you can expect:
- 初始快照式复制: By default, Reindex uses a scroll search under the hood, which takes a point-in-time snapshot of your source index the moment the job starts. This means all documents that existed in the source index at that exact time will be copied to the target index.
- 新增/更新文档不会自动同步: Any documents added, updated, or deleted in the source index after the Reindex job starts won’t be included in the initial run. The job doesn’t actively listen for changes to the source index during execution.
- 冲突处理: If a document in the source index gets updated while Reindex is copying it, you might hit version conflicts. You can handle this by adding
"conflicts": "proceed"to your request to skip conflicting documents, or use versioning settings to prioritize the latest version.
Example of a basic Reindex request with conflict handling:
POST _reindex { "conflicts": "proceed", "source": { "index": "your_source_index" }, "dest": { "index": "your_target_index" } }
二、Reindex任务会持续执行吗?
Short answer: No, it won’t run indefinitely by default.
- The Reindex job will finish once it has copied all documents from the initial snapshot. It doesn’t automatically restart or continue syncing new writes to the source index.
- If you need to keep the target index in sync with ongoing writes to the source, you have two main options:
- Incremental Reindex runs: Set up a scheduled task (like a cron job or Airflow DAG) that runs periodic Reindex requests targeting only new/updated documents. You can filter using a timestamp field (like
@timestamp) or the document’s_versionnumber to avoid reprocessing old data. Here’s an example:POST _reindex { "source": { "index": "your_source_index", "query": { "range": { "@timestamp": { "gte": "2024-05-20T00:00:00Z" // Timestamp of your last sync } } } }, "dest": { "index": "your_target_index", "op_type": "create" // Prevents duplicates for already synced docs } } - Cross-Cluster Replication (CCR): If you’re using Elasticsearch 6.7+, CCR is built for continuous, near-real-time syncing between indices (even across clusters). It’s a more robust solution than scheduled Reindex jobs for ongoing synchronization.
- Incremental Reindex runs: Set up a scheduled task (like a cron job or Airflow DAG) that runs periodic Reindex requests targeting only new/updated documents. You can filter using a timestamp field (like
Key Notes for Production
- Always test Reindex jobs in a staging environment first to mimic production write loads.
- Monitor the Reindex task using
GET _tasks?detailed=true&actions=*reindexto track progress and spot issues early. - For large indices, use sliced Reindex (
"slice": {"size": 5}) to parallelize the job and reduce runtime.
内容的提问来源于stack exchange,提问作者luca.p.alexandru
相关产品推荐
相关产品推荐

