Kubernetes从1.23升级至1.25后Fluentd出现Pattern不匹配错误
Kubernetes 1.25升级后Fluentd日志解析异常问题
问题现象
原在Kubernetes 1.23环境中运行正常的Fluentd日志采集配置,升级至1.25版本后出现大量日志解析警告,示例警告内容如下:
2023-04-03 01:32:06 +0000 [warn]: #0 [in_tail_container_logs] pattern not matched: "2023-04-03T01:32:02.9256618Z stdout F [2023-04-03T01:32:02.925Z] DEBUG transaction-677fffdfc4-tc4rx-18/TRANSPORTER: NATS client pingTimer: 1"
当前使用的配置文件
container.conf 配置
<source> @type tail @id in_tail_container_logs @label @containers path /var/log/containers/*.log exclude_path ["/var/log/containers/fluentd*"] pos_file /var/log/fluentd-containers.log.pos tag * read_from_head true <parse> @type json time_format %Y-%m-%dT%H:%M:%S.%NZ </parse> </source> <label @containers> <filter **> @type kubernetes_metadata @id filter_kube_metadata </filter> <filter **> @type record_transformer @id filter_containers_stream_transformer <record> stream_name ${tag_parts[3]} </record> </filter> <match **> @type cloudwatch_logs @id out_cloudwatch_logs_containers region "#{ENV.fetch('AWS_REGION')}" log_group_name "/k8s-nest/#{ENV.fetch('AWS_EKS_CLUSTER_NAME')}/containers" log_stream_name_key stream_name remove_log_stream_name_key true auto_create_stream true <buffer> flush_interval 5 chunk_limit_size 2m queued_chunks_limit_size 32 retry_forever true </buffer> </match> </label>
Fluentd DaemonSet 配置
apiVersion: apps/v1 kind: DaemonSet metadata: labels: k8s-app: fluentd-cloudwatch name: fluentd-cloudwatch namespace: kube-system spec: selector: matchLabels: k8s-app: fluentd-cloudwatch template: metadata: labels: k8s-app: fluentd-cloudwatch annotations: iam.amazonaws.com/role: fluentd spec: serviceAccount: fluentd serviceAccountName: fluentd containers: - env: - name: AWS_REGION value: us-west-1 - name: AWS_EKS_CLUSTER_NAME value: dex-eks-west #image: 'fluent/fluentd-kubernetes-daemonset:v1.1-debian-cloudwatch' image: 'fluent/fluentd-kubernetes-daemonset:v1.15.3-debian-cloudwatch-1.1' imagePullPolicy: IfNotPresent name: fluentd-cloudwatch resources: limits: memory: 200Mi requests: cpu: 100m memory: 200Mi terminationMessagePath: /dev/termination-log terminationMessagePolicy: File volumeMounts: - mountPath: /config-volume name: config-volume - mountPath: /fluentd/etc name: fluentdconf - mountPath: /var/log name: varlog - mountPath: /var/lib/docker/containers name: varlibdockercontainers readOnly: true - mountPath: /run/log/journal name: runlogjournal readOnly: true dnsPolicy: ClusterFirst initContainers: - command: - sh - '-c' - cp /config-volume/..data/* /fluentd/etc image: busybox imagePullPolicy: Always name: copy-fluentd-config resources: {} terminationMessagePath: /dev/termination-log terminationMessagePolicy: File volumeMounts: - mountPath: /config-volume name: config-volume - mountPath: /fluentd/etc name: fluentdconf terminationGracePeriodSeconds: 30 volumes: - configMap: defaultMode: 420 name: fluentd-config name: config-volume - emptyDir: {} name: fluentdconf - hostPath: path: /var/log type: '' name: varlog - hostPath: path: /var/lib/docker/containers type: '' name: varlibdockercontainers - hostPath: path: /run/log/journal type: '' name: runlogjournal
问题原因
Kubernetes 1.24及后续版本,默认容器日志格式从JSON格式切换为CRI日志格式(即时间戳 流类型 标志 日志内容的纯文本格式)。当前Fluentd配置仍使用@type json解析日志,无法匹配新的CRI格式,因此触发pattern not matched警告。
解决方案
核心修改:替换日志解析器
修改container.conf中<source>部分的<parse>配置,使用CRI专用解析器:
<source> @type tail @id in_tail_container_logs @label @containers path /var/log/containers/*.log exclude_path ["/var/log/containers/fluentd*"] pos_file /var/log/fluentd-containers.log.pos tag * read_from_head true <parse> @type cri time_format %Y-%m-%dT%H:%M:%S.%N%:z </parse> </source>
关键说明
@type cri:专门适配CRI格式的日志解析器,可自动识别时间戳、stdout/stderr流类型、日志内容等字段time_format %Y-%m-%dT%H:%M:%S.%N%:z:匹配CRI日志的时间戳格式(如示例中的2023-04-03T01:32:02.9256618Z)
兼容旧日志格式(可选)
如果集群中存在部分旧容器仍输出JSON格式日志,可使用multi_format同时兼容两种格式:
<parse> @type multi_format <pattern> @type cri time_format %Y-%m-%dT%H:%M:%S.%N%:z </pattern> <pattern> @type json time_format %Y-%m-%dT%H:%M:%S.%NZ </pattern> </parse>
内容的提问来源于stack exchange,提问作者Vijayalakshmi Natarajan
相关产品推荐
相关产品推荐

