无法将索引字段值传入Painless脚本的问题咨询
问题场景
索引结构
PUT my-index-000001/_doc/1 { "virtual": "/testss/3-1.pdf", "file": "3-1", "caseno": "testss" }
尝试的搜索脚本
想要通过Painless脚本分割file字段的值,根据分割后的长度计算新字段:
GET my-index-000001/_search { "script_fields": { "mynewfield": { "script": { "source":""" List i=Arrays.asList(doc['file'].value.splitOnToken("-")); if (i.length==1){ return Float.parseFloat(i[0]); } if (i.length==2){ return Float.parseFloat(i[0])+Float.parseFloat(i[1])/100; } """ } } } }
报错核心信息
执行后返回400错误,关键原因:
Text fields are not optimised for operations that require per-document field data like aggregations and sorting, so these operations are disabled by default. Please use a keyword field instead. Alternatively, set fielddata=true on [file] in order to load field data by uninverting the inverted index. Note that this can use significant memory.
错误原因
file字段默认是text类型,这类字段主打全文检索能力,默认不启用字段数据(fielddata)存储;而通过doc['file'].value访问字段需要加载字段数据,因此触发报错。
解决方案
方案一:直接通过_source获取原始值(临时场景适用)
无需修改索引映射,直接从原始文档中读取file字段的原始字符串值:
GET my-index-000001/_search { "script_fields": { "mynewfield": { "script": { "source":""" String fileValue = params._source.file; List i=Arrays.asList(fileValue.splitOnToken("-")); if (i.length==1){ return Float.parseFloat(i[0]); } if (i.length==2){ return Float.parseFloat(i[0])+Float.parseFloat(i[1])/100; } """ } } } }
方案二:添加keyword子字段(推荐长期使用)
给file字段新增keyword类型的子字段,这类字段专门用于精确匹配、脚本操作和聚合,性能更稳定:
- 更新索引映射:
PUT my-index-000001/_mapping { "properties": { "file": { "type": "text", "fields": { "keyword": { "type": "keyword" } } } } }
- 修改脚本访问
file.keyword字段:
GET my-index-000001/_search { "script_fields": { "mynewfield": { "script": { "source":""" List i=Arrays.asList(doc['file.keyword'].value.splitOnToken("-")); if (i.length==1){ return Float.parseFloat(i[0]); } if (i.length==2){ return Float.parseFloat(i[0])+Float.parseFloat(i[1])/100; } """ } } } }
不推荐的方案:开启fielddata
虽然报错提示可以开启fielddata=true,但text字段的fielddata会在JVM堆内存中加载大量倒排索引数据,极易导致内存溢出或性能下降,因此不建议使用。
内容的提问来源于stack exchange,提问作者ScottyCov

