解决Fluentd向OpenSearch传输日志时的Mapper Parsing Exception问题
问题分析
报错根源是日志中的\x92属于Windows-1252编码的右单引号(ASCII 146),当被当作UTF-8解析时属于无效字符序列,OpenSearch默认会拒绝包含这类无效UTF-8字符的文档,从而抛出mapper_parsing_exception。
解决方法
方法1:在Fluentd中替换无效字符
在Fluentd的source和match配置之间添加mutate过滤器,直接将\x92替换为合法的单引号:
<filter Connectship.**> @type mutate <replace> field log pattern /\x92/ replacement "'" </replace> </filter>
方法2:在Fluent-bit中用Lua过滤器处理(比modify更灵活)
之前的modify过滤器用法错误(它用于重命名字段,而非修改字段内容),改用Lua过滤器替换字段内容:
- 添加Fluent-bit过滤器配置:
[FILTER] Name lua Match * script replace_invalid_chars.lua call replace_log_chars
- 创建
replace_invalid_chars.lua脚本:
function replace_log_chars(tag, timestamp, record) if record["log"] then record["log"] = string.gsub(record["log"], "\x92", "'") end return 2, timestamp, record end
方法3:修改OpenSearch索引映射忽略无效字符
如果不想修改日志内容,可以给log字段设置ignore_malformed: true,让OpenSearch跳过无效字符而非拒绝文档:
- 创建索引模板(若索引已存在,需先删除重建或使用reindex):
PUT _index_template/connectship_template { "index_patterns": ["connectship11*"], "template": { "mappings": { "properties": { "log": { "type": "text", "ignore_malformed": true } } } } }
方法4:在Fluentd中转换日志编码
如果日志本身是Windows-1252编码,用detect_encoding过滤器将其转为UTF-8,\x92会被正确解析为UTF-8格式的右单引号:
<filter Connectship.**> @type detect_encoding field log target_encoding utf-8 fallback_encoding windows-1252 </filter>
内容的提问来源于stack exchange,提问作者AmythD
相关产品推荐
相关产品推荐

