K8s集群EFK栈中如何配置Fluentd仅采集指定多容器的日志
Problem Statement
I've already set up the EFK stack in my Kubernetes cluster. Currently, Fluentd is collecting logs from all containers in the cluster, but I want to adjust the configuration to only collect logs from containers A, B, C, and D. When containers have a unified prefix (like A-app), I can use a pattern like *-app.log, but these target containers don't share a common prefix.
Here's my current fluentd-inputs.conf:
# HTTP input for the liveness and readiness probes <source> @type http port 9880 </source> # Get the logs from the containers running in the node <source> @type tail path /var/log/containers/*-app.log // How to configure this to match multiple different containers? # exclude Fluentd logs exclude_path /var/log/containers/*fluentd*.log pos_file /opt/bitnami/fluentd/logs/buffers/fluentd-docker.pos tag kubernetes.* read_from_head true <parse> @type json </parse> </source> # enrich with kubernetes metadata <filter kubernetes.**> @type kubernetes_metadata </filter>
How should I modify the path parameter in the above configuration to accurately collect logs only from containers A, B, C, and D?
Answer
Great question! When targeting specific containers without a shared prefix, you have a few straightforward options to tweak your Fluentd configuration. Let's go through the most effective approaches:
Option 1: Use Multiple Glob Patterns
Kubernetes container log files in /var/log/containers/ follow the format <container-name>_<namespace>_<pod-name>_<container-id>.log. You can directly target each container with a glob pattern that matches its name at the start of the filename, separating multiple patterns with commas:
<source> @type tail # Match logs from containers A, B, C, D explicitly path /var/log/containers/A_*.log,/var/log/containers/B_*.log,/var/log/containers/C_*.log,/var/log/containers/D_*.log # exclude Fluentd logs exclude_path /var/log/containers/*fluentd*.log pos_file /opt/bitnami/fluentd/logs/buffers/fluentd-docker.pos tag kubernetes.* read_from_head true <parse> @type json </parse> </source>
This is simple to read and maintain, especially if you only have a small set of containers to track.
Option 2: Use a Regular Expression for Conciseness
If you want a more compact configuration, use a regex to match any of your target container names in the filename. Fluentd's tail source supports regex-based path matching directly in the path parameter:
<source> @type tail # Regex to match any of containers A, B, C, D at the start of the log file path /var/log/containers/(A|B|C|D)_*.log # exclude Fluentd logs exclude_path /var/log/containers/*fluentd*.log pos_file /opt/bitnami/fluentd/logs/buffers/fluentd-docker.pos tag kubernetes.* read_from_head true <parse> @type json </parse> </source>
The (A|B|C|D) group matches any of your container names, followed by the rest of the standard Kubernetes log filename structure.
Option 3: Filter After Collection (For Extra Flexibility)
If you need more control—like targeting containers by namespace, labels, or if container names might change—you can collect all logs first, then filter out everything except your target containers using the kubernetes_metadata filter data:
# Keep collecting all container logs initially <source> @type tail path /var/log/containers/*.log exclude_path /var/log/containers/*fluentd*.log pos_file /opt/bitnami/fluentd/logs/buffers/fluentd-docker.pos tag kubernetes.* read_from_head true <parse> @type json </parse> </source> # Enrich logs with Kubernetes metadata (including container_name) <filter kubernetes.**> @type kubernetes_metadata </filter> # Filter to retain only logs from containers A, B, C, D <filter kubernetes.**> @type grep <regexp> key $.kubernetes.container_name pattern ^(A|B|C|D)$ </regexp> </filter>
This method uses the grep filter to check the container_name field (populated by the metadata filter) and only keeps logs from your specified containers. It's ideal if you need to base filtering on more than just the filename.
内容的提问来源于stack exchange,提问作者confusedWarrior

