ElasticSearch rank_eval报错:Failed to build [request] after last required field arrived
问题描述
在LinkSO数据集上使用ElasticSearch开展排序评估工作时,调用es.rank_eval()函数触发400错误:
BadRequestError: BadRequestError(400, 'x_content_parse_exception', 'Failed to build [request] after last required field arrived')
相关代码及配置如下:
1. 获取NDCG的get_ndcg函数
def get_ndcg(dataframe, input_index="java_lm"): ndcg_list = [] # loop through each qid1 for i in range(0, len(dataframe["qid1"]), 30): qid1_title = es.get(index=input_index, id=dataframe["qid1"][i])['_source']['title'] # load ratings from the json file f = open("qids/" + input_index + "/" + str(dataframe["qid1"][i]) + ".json") data = json.load(f) _search = ranking(dataframe["qid1"][i], qid1_title, ratings=data) result = es.rank_eval(index=input_index, body=_search) ndcg = result['metric_score'] ndcg_list.append(ndcg) return ndcg_list
2. 生成排序请求的ranking函数
def ranking(qid1, qid1_title, ratings): _search = { "requests": [ { "id": str(qid1), "request": { "query": { "bool": { "must_not": { "match": { "_id": qid1 } }, "should": [ { "match": { "title": { "query": qid1_title, "boost": 3.0, "analyzer": "my_analyzer" } } }, { "match": { "body": { "query": qid1_title, "boost": 0.5, "analyzer": "my_analyzer" } } }, { "match": { "answer": { "query": qid1_title, "boost": 0.5, "analyzer": "my_analyzer" } } } ] } } }, "ratings": ratings } ], "metric": { "dcg": { "k": 10, "normalize": True } } } return _search
3. ratings文件格式
[ {"_index": "java_lm", "_id": "15194804", "rating": 0}, {"_index": "java_lm", "_id": "18264178", "rating": 0}, {"_index": "java_lm", "_id": "16225177", "rating": 1}, {"_index": "java_lm", "_id": "16445238", "rating": 0}, {"_index": "java_lm", "_id": "17233226", "rating": 0} ]
4. Kibana创建索引的PUT命令
PUT /java_lm { "settings": { "similarity": { "LM": { "type": "LMDirichlet", "mu": 2000 } }, "analysis": { "analyzer": { "my_analyzer": { "tokenizer": "whitespace", "filter": [ "lowercase", "porter_stem" ] } } } }, "mappings": { "properties": { "title": { "type": "text", "analyzer": "my_analyzer", "similarity": "LM" }, "body": { "type": "text", "analyzer": "my_analyzer", "similarity": "LM" }, "answer": { "type": "text", "analyzer": "my_analyzer", "similarity": "LM" } } } }
问题排查与修正方案
核心错误原因
- Metric字段格式不兼容:使用
dcg类型并指定normalize: true的写法在部分Elasticsearch版本中存在解析问题,官方更推荐直接使用ndcg类型(默认自动归一化)。 - _id查询方式错误:
_id是关键字字段,match查询会对输入分词,导致无法精准匹配,应该用term查询。 - 文件读取未规范关闭:直接
open文件后未关闭,可能导致资源泄漏或读取异常。
具体修正步骤
1. 修正Metric配置
将metric部分改为ndcg类型,无需额外指定归一化参数:
"metric": { "ndcg": { "k": 10 } }
2. 修正must_not子句的查询方式
用term查询替代match查询_id,确保精准匹配:
"must_not": { "term": { "_id": str(qid1) # 强制转为字符串,避免类型不匹配 } }
3. 规范文件读取方式
使用上下文管理器with自动关闭文件,避免资源问题:
with open(f"qids/{input_index}/{str(dataframe['qid1'][i])}.json", "r") as f: data = json.load(f)
4. 完整修正后的ranking函数
def ranking(qid1, qid1_title, ratings): _search = { "requests": [ { "id": str(qid1), "request": { "query": { "bool": { "must_not": { "term": { "_id": str(qid1) } }, "should": [ { "match": { "title": { "query": qid1_title, "boost": 3.0, "analyzer": "my_analyzer" } } }, { "match": { "body": { "query": qid1_title, "boost": 0.5, "analyzer": "my_analyzer" } } }, { "match": { "answer": { "query": qid1_title, "boost": 0.5, "analyzer": "my_analyzer" } } } ] } } }, "ratings": ratings } ], "metric": { "ndcg": { "k": 10 } } } return _search
5. 额外优化:循环逻辑合理性
原循环步长30可能跳过部分qid1,建议改为遍历所有唯一qid1:
for qid1 in dataframe["qid1"].unique(): qid1_title = es.get(index=input_index, id=qid1)['_source']['title'] # 后续读取文件、生成请求逻辑不变
内容的提问来源于stack exchange,提问作者insanely_a_
相关产品推荐
相关产品推荐

