Elasticsearch重新索引文档丢失,无报错求排查方案
首先看你的reindex任务详情,total:2414,updated:1346+created:1068=2414,说明任务本身已经处理了所有源文档,不存在任务层面的文档丢失。你觉得“丢失”大概率是以下原因之一,按步骤排查:
检查失败文档索引
你的摄入管道配置了on_failure逻辑,会把处理失败的文档写入failed-collection-with-embeddings索引。执行以下命令查看该索引的文档数:GET failed-collection-with-embeddings/_count如果有文档,查看具体失败原因:
GET failed-collection-with-embeddings/_search { "query": { "match_all": {} }, "_source": ["ingest.failure"] }验证目标索引实际文档数
不要用普通搜索结果判断文档数量,直接查询目标索引的总计数:GET collection-with-embeddings/_count如果计数是2414,说明文档都在,只是你搜索时的查询条件、分页设置或者索引未刷新导致没看到全部。可以手动刷新索引后再查:
POST collection-with-embeddings/_refresh排查摄入管道的字段映射问题
你的inference处理器配置了field_map: {"text": "text_field"},意思是把源文档的text字段作为模型输入的text_field。如果源文档中没有text字段,或者字段名不匹配(比如是content、title等),模型会处理失败,文档被转去失败索引。
可以先随机查看几个源文档的字段结构:GET collection/_search { "size": 5 }如果源文档没有
text字段,需要修改field_map为实际的字段名,比如源字段是content,就改成:"field_map": { "content": "text_field" }检查模型输入长度限制
sentence-transformers__msmarco-minilm-l-12-v3模型有输入文本长度限制(通常是512个token),如果源文档的目标字段内容过长,会触发处理失败。查看失败文档的ingest.failure字段就能确认这个问题。如果是这个原因,可以在inference处理器前添加truncate处理器截断文本:PUT _ingest/pipeline/text-embeddings { "description": "Text embedding pipeline", "processors": [ { "truncate": { "field": "text", "length": 512, "preserve_position": true } }, { "inference": { "model_id": "sentence-transformers__msmarco-minilm-l-12-v3", "target_field": "text_embedding", "field_map": { "text": "text_field" } } } ], "on_failure": [ { "set": { "description": "Index document to 'failed-<index>'", "field": "_index", "value": "failed-{{{_index}}}" } }, { "set": { "description": "Set error message", "field": "ingest.failure", "value": "{{_ingest.on_failure_message}}" } } ] }
内容的提问来源于stack exchange,提问作者Sazzad

