Fluentd异常将容器日志输出至自身STDOUT问题排查求助
环境信息
- GKE集群版本:1.21.12-gke.1700
- Fluentd Kubernetes Daemonset版本:v1.14.3(镜像:
fluent/fluentd-kubernetes-daemonset:v1.14.3-debian-gcs-1.1) - 核心目标:采集指定容器日志并转发至GCS Bucket
异常现象
部分容器的日志会偶发直接输出到Fluentd自身的STDOUT中,而非通过fluent-plugin-gcs插件上传至GCS。典型表现如下:
2022-07-20 13:29:04 +0000 [info]: #0 [in_tail_container_logs] following tail of /var/log/containers/my-pod_my-namespace_helloworld-7e6359514a5601e5ad1823d145fd3b73f7b65648f5cb760f2c1855dabe27d606.log"
...
HELLO WORLD
...
容器日志本身是JSON结构化格式,但此处直接以明文形式出现在Fluentd的运行日志里。
已排查项
- 开启
fluentd -vv高 verbose 模式,未发现相关错误线索 - 问题随机出现,仅部分容器偶发异常,部分容器从未出现该问题
- 异常容器无特殊配置或运行特征
- 确认未配置任何Fluentd输出插件将日志发送至STDOUT
当前Fluentd配置
<label @FLUENT_LOG> <match fluent.**> @type null </match> </label> <source> @type tail @id in_tail_container_logs path /var/log/containers/*.log exclude_path ["/var/log/containers/fluentd-*", "/var/log/containers/fluentbit-*", "/var/log/containers/kube-*", "/var/log/containers/pdsci-*", "/var/log/containers/gke-*"] pos_file /var/log/fluentd-containers.log.pos tag "kubernetes.*" refresh_interval 1s read_from_head true follow_inodes true <parse> @type json time_format %Y-%m-%dT%H:%M:%S.%NZ keep_time_key true </parse> </source> <filter kubernetes.**> @type kubernetes_metadata @id filter_kube_metadata kubernetes_url "#{'https://' + ENV.fetch('KUBERNETES_SERVICE_HOST') + ':' + ENV.fetch('KUBERNETES_SERVICE_PORT') + '/api'}" verify_ssl true ca_file "#{ENV['KUBERNETES_CA_FILE']}" watch false # Don't watch for changes in container metadata de_dot false # Don't replace dots in labels and annotations skip_labels false skip_master_url true skip_namespace_metadata true annotation_match ["app.*\/log-.+"] </filter> <filter kubernetes.**> @type grep <regexp> key $.kubernetes.namespace_name pattern /^my-namespace$/ </regexp> <regexp> key $['kubernetes']['labels']['example.com/collect'] pattern /^yes$/ </regexp> </filter> <match kubernetes.**> @type gcs @id out_gcs project "#{ENV['GCS_BUCKET_PROJECT']}" bucket "#{ENV.fetch('GCS_BUCKET_PROJECT')" object_key_format %Y%m%d/%H%M/%{$.kubernetes.pod_name}_${$.kubernetes.container_name}_${$.docker.container_id}/%{index}.%{file_extension} store_as json <buffer time,$.kubernetes.pod_name,$.kubernetes.container_name,$.docker.container_id> @type file path /var/log/fluentd-buffers/gcs.buffer timekey 30 timekey_wait 5 timekey_use_utc true # use utc chunk_limit_size 1MB flush_at_shutdown true </buffer> <format> @type json </format> </match>
排查与修复方向
修复GCS插件配置语法错误
配置中bucket字段存在语法错误:bucket "#{ENV.fetch('GCS_BUCKET_PROJECT')"缺少闭合的}。该错误会导致GCS插件初始化失败,此时Fluentd会将无法匹配到有效输出插件的日志默认输出到STDOUT。修复后应为:bucket "#{ENV.fetch('GCS_BUCKET_PROJECT')}"捕获JSON解析失败的日志
Fluentd的in_tail插件若遇到无法解析的JSON行,会默认将原始文本输出到STDOUT。可以在source段添加错误路由,将解析失败的日志写入文件排查:<source> <!-- 原有source配置 --> @label @ERROR </source> <label @ERROR> <match **> @type file path /var/log/fluentd-parse-errors.log format json_lines </match> </label>排查日志过滤路由异常
在grep过滤器后添加临时输出,捕获未被过滤的日志,确认是否存在元数据注入失败导致的路由异常:<filter kubernetes.**> <!-- 原有grep配置 --> </filter> <match kubernetes.**> @type file path /var/log/fluentd-unmatched.log format json_lines @id out_unmatched </match>检查缓冲目录权限
确认Fluentd进程对/var/log/fluentd-buffers/gcs.buffer目录有读写权限,权限不足会导致GCS插件缓冲失败,触发默认STDOUT输出。
内容的提问来源于stack exchange,提问作者ButterDog

