Fluentd同步Elasticsearch遇数组类型冲突报错求助
问题:Fluentd写入Elasticsearch因嵌套数组类型冲突报错400
问题背景
原始日志结构如下:
{ "key": "value", "inputs": [ [ "2023-01-16T04:45:12.238Z", { "type": "channel", "subtype": "profile", "data": { "firstName": "Customer" } } ] ] }
写入Elasticsearch时触发400错误:
#0 dump an error event: error_class=Fluent::Plugin::ElasticsearchErrorHandler::ElasticsearchError error="400 - Rejected by Elasticsearch [error type]: illegal_argument_exception [reason]: 'can't merge a non object mapping [data.inputs] with an object mapping'" location=nil
核心原因是inputs数组的子元素为混合类型(字符串日期+对象),Elasticsearch无法为该字段生成统一的映射规则。
解决方法
方法1:重构日志结构(推荐)
使用Fluentd的record_transformer插件,将子数组中的日期和对象拆分为结构化字段,让inputs成为纯对象数组,既保留全部数据又符合ES映射要求。
在Fluentd配置文件中添加以下过滤规则:
<filter **> @type record_transformer enable_ruby true <record> inputs ${record['inputs'].map { |arr| { 'timestamp' => arr[0], 'event' => arr[1] } }} </record> </filter>
处理后日志结构变为:
{ "key": "value", "inputs": [ { "timestamp": "2023-01-16T04:45:12.238Z", "event": { "type": "channel", "subtype": "profile", "data": { "firstName": "Customer" } } } ] }
方法2:移除日期字段(若无需保留)
如果日志中的日期信息无用,可直接过滤掉子数组中的日期元素,仅保留对象部分:
<filter **> @type record_transformer enable_ruby true <record> inputs ${record['inputs'].map { |arr| arr[1] }} </record> </filter>
处理后inputs数组仅包含对象,可正常写入Elasticsearch。
方法3:手动创建ES索引映射
提前为目标索引创建自定义映射,指定inputs为嵌套类型(灵活性较差,仅适合固定日志格式场景):
PUT /your-index-name { "mappings": { "properties": { "inputs": { "type": "nested", "properties": { "0": {"type": "date"}, "1": {"type": "object"} } } } } }
内容的提问来源于stack exchange,提问作者Aayush Raj
相关产品推荐
相关产品推荐

