Elasticsearch中@message.keyword存在但size为0的问题排查求助
问题描述
有一条Elasticsearch日志条目,其@message字段内容如下:
14-Apr-2023 20:44:46.693 INFO [pool-2-thread-24] com.xyz.log [app_id:uuid] calling execute-task with url=https://example.com/api/applications/uuid/tasks/TASK_NAME/execute/, body=workflow_id=uuid variables={
使用以下脚本生成脚本字段时,该条目返回doc[@message.keyword].size()==0,但部分短日志(如下方示例)能正确提取出线程名(如Thread-4):
if(doc.containsKey('@message') ) { if (doc.containsKey('@message.keyword')) { if(doc['@message.keyword'].size()==0){ return "doc[@message.keyword].size()==0" } else { String message= doc['@message.keyword'].value; int start=message.indexOf('['); int end=message.indexOf(']'); String x= message.substring(start+1,end); return x; } } else { return "!doc.containsKey('@message.keyword')"; } } else { return "!doc.containsKey('@message')"; }
可正常提取的短日志示例:
19-Apr-2023 15:10:28.113 FINE [Thread-4] org.camunda.commons.logging.BaseLogger.logDebug ENGINE-13011 closing existing command context
自动生成的索引映射如下:
{ "dpprod-workflow-web.2023.04" : { "mappings" : { "properties" : { "@id" : { "type" : "text", "fields" : { "keyword" : { "type" : "keyword", "ignore_above" : 256 } } }, "@log_group" : { "type" : "text", "fields" : { "keyword" : { "type" : "keyword", "ignore_above" : 256 } } }, "@log_stream" : { "type" : "text", "fields" : { "keyword" : { "type" : "keyword", "ignore_above" : 256 } } }, "@message" : { "type" : "text", "fields" : { "keyword" : { "type" : "keyword", "ignore_above" : 256 } } }, "@owner" : { "type" : "text", "fields" : { "keyword" : { "type" : "keyword", "ignore_above" : 256 } } }, "@timestamp" : { "type" : "date" } } } } }
另外尝试开启field_data=true后使用@message字段,得到的结果是0a9b7ed24c9b12而非实际字符串。
原因分析
@message.keyword无值的原因
映射中@message.keyword的ignore_above设置为256,意味着当@message的原始字符串长度超过256个字符时,Elasticsearch不会生成该字段的keyword值。此时doc.containsKey('@message.keyword')返回true(因为字段在映射中存在),但doc['@message.keyword']没有实际内容,所以size()返回0。你提供的长日志明显超过了256字符限制,因此触发了这个规则。开启field_data后得到哈希值的原因
text类型字段开启field_data=true后,Elasticsearch会为分词后的词项生成哈希映射用于聚合等操作,脚本中直接获取的value是词项的哈希值,而非原始字符串,因此得到的是类似0a9b7ed24c9b12的结果,这并不是你需要的原始日志内容。
解决方案
方案1:调整keyword字段的长度限制
修改@message.keyword的ignore_above参数为更大的值(比如1024,根据实际日志最大长度调整),注意keyword字段的最大支持长度为32766字节。修改映射的示例请求:
PUT /dpprod-workflow-web.2023.04/_mapping { "properties": { "@message": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 1024 } } } } }
修改后需要重新索引数据,新的日志条目会生成完整的keyword值,旧数据需要重新索引才能生效。
方案2:直接从_source获取原始日志
脚本中直接使用params._source['@message']获取完整的原始日志内容,不受分词或ignore_above限制,这是更可靠的方式。修改后的脚本如下:
if (params._source?.@message) { String message = params._source['@message']; int start = message.indexOf('['); int end = message.indexOf(']', start); // 指定从start之后找第一个],避免匹配到[app_id:uuid]里的] if (start != -1 && end != -1) { return message.substring(start + 1, end); } else { return "No thread name found"; } } else { return "@message field missing"; }
这个脚本直接读取原始日志内容,无需依赖keyword字段,也不会受长度限制影响,同时优化了索引查找逻辑,避免匹配到日志中其他方括号的内容。
内容的提问来源于stack exchange,提问作者Pramod Jangam

