You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Otel Collector Filelog接收器日志轮转丢事件及重复读取问题咨询

exporters:
splunk_hec/audit:
token: ${SPLUNK_HEC_TOKEN}
endpoint: https://my.splunk.forwarder:8088/services/collector
disable_compression: false
tls:
insecure_skip_verify: false
ca_file: /etc/splunk/certs/hec_ca_file
idle_conn_timeout: 10s
profiling_data_enabled: false
retry_on_failure:
enabled: true
initial_interval: 5s
max_elapsed_time: 300s
max_interval: 30s
sending_queue:
enabled: true
num_consumers: 10
queue_size: 5000
extensions:
file_storage:
directory: /var/addon/splunk/otel_pos
health_check:
endpoint: 0.0.0.0:13133
processors:
batch/audit:
timeout: 5s
send_batch_size: 2048
send_batch_max_size: 8192
filter/audit_logs:
logs:
log_record:
- IsMatch(ParseJSON(body)["verb"], "get") and not IsMatch(ParseJSON(body)["requestURI"], "./secrets/.$")
- IsMatch(ParseJSON(body)["verb"], "update") and IsMatch(ParseJSON(body)["requestURI"], "/apis/conjur.cyberark.com/v1/namespaces/tenant-conjur-follower./conjurfollowers/conjur-follower./status$")
- IsMatch(ParseJSON(body)["verb"], "patch") and IsMatch(ParseJSON(body)["requestURI"], "/apis/conjur.cyberark.com/v1/namespaces/tenant-conjur-follower./conjurfollowers/conjur-follower./status$")

- IsMatch(ParseJSON(body)["verb"], "list")
   - IsMatch(ParseJSON(body)["verb"], "watch")

resource/add_log_metadata:
attributes:
- action: insert
key: com.splunk.index
value: gen_ocp_pa
- action: insert
key: k8s.cluster.name
value: ${K8S_CLUSTERNAME}
- action: insert
key: log_type
value: audit
- action: insert
key: host.name
value: ${K8S_NODE_NAME}
memory_limiter:
check_interval: 2s
limit_percentage: 80
resourcedetection:
detectors:
- env
- system
override: true
timeout: 15s
receivers:
filelog/openshift-audit:
include:
- /var/log/oauth-apiserver/audit.log
- /var/log/kube-apiserver/audit.log
- /var/log/openshift-apiserver/audit.log
start_at: end
storage: file_storage
retry_on_failure:
enabled: true
include_file_path: true
include_file_name: true
poll_interval: 1ms
operators:
- type: json_parser
id: json_parser
parse_from: body
on_error: send
timestamp:
parse_from: attributes.requestReceivedTimestamp
layout: '%Y-%m-%dT%H:%M:%S.%f%z'
- type: copy
from: attributes['log.file.path']
to: attributes['com.splunk.source']
nop:
service:
extensions:

  • file_storage
  • health_check
    pipelines:
    logs/audit:
    receivers:
    • filelog/openshift-audit
      processors:
    • filter/audit_logs
    • resource/add_log_metadata
    • memory_limiter
    • batch/audit
      exporters:
    • splunk_hec/audit
      telemetry:
      metrics:
      address: 0.0.0.0:8890
      logs:
      level: info
# 解决方案
针对日志轮转引发的事件丢失与重复读取问题,可通过以下配置调整和优化手段解决:

## 1. 正确配置轮转日志匹配规则,避免重复读取
修改filelog接收器的`include`规则,用通配符覆盖所有轮转后的日志文件,同时排除已压缩的归档文件,配合`file_storage`的位置跟踪功能(通过inode识别文件),确保不会重复读取已处理内容:
```yaml
receivers:
  filelog/openshift-audit:
    include:
      - /var/log/oauth-apiserver/audit.log*
      - /var/log/kube-apiserver/audit.log*
      - /var/log/openshift-apiserver/audit.log*
    exclude:
      - /var/log/oauth-apiserver/audit.log.*.gz
      - /var/log/kube-apiserver/audit.log.*.gz
      - /var/log/openshift-apiserver/audit.log.*.gz
    # 保留原有其他配置
  • 通配符*覆盖轮转后的audit.log.1、audit.log.2等未压缩文件
  • 排除.gz压缩文件,这类文件不会再写入新内容,无需持续监控
  • file_storage会持久化每个文件的inode和读取偏移量,即使文件重命名,Collector也能准确识别并从上次位置继续读取

2. 优化JSON解析与事件完整性保障

当前json_parser的on_error: send会直接发送解析失败的不完整事件,可调整错误策略并添加多行合并规则,确保读取完整的JSON事件:

operators:
  # 新增多行合并规则,确保完整读取JSON对象
  - type: multiline
    firstline: '^{'
    max_lines: 100 # 根据实际审计日志的最大行数调整
    timeout: 500ms
  - type: json_parser
    id: json_parser
    parse_from: body
    on_error: drop # 或改为"retry",配合重试机制
    timestamp:
      parse_from: attributes.requestReceivedTimestamp
      layout: '%Y-%m-%dT%H:%M:%S.%f%z'
  # 保留原有copy操作符
  • multiline操作符会等待以{开头的完整JSON对象,避免读取日志轮转时截断的不完整行
  • 若选择on_error: drop,会丢弃无法解析的不完整事件;若需重试,可设置on_error: retry并配合接收器的retry_on_failure配置,等待事件完整后再处理

3. 调整日志读取的轮询与缓冲策略

当前poll_interval: 1ms过于频繁,易导致轮转时读取截断内容,可调整为更合理的参数:

receivers:
  filelog/openshift-audit:
    poll_interval: 100ms # 降低轮询频率,减少文件系统查询
    read_buffer_size: 65536 # 增大读取缓冲区,确保一次性读取完整事件
    # 保留原有其他配置
  • 适当增大poll_interval可降低文件系统的负载,减少轮转瞬间读取不完整内容的概率
  • 增大read_buffer_size有助于一次性读取较大的日志行,避免拆分事件

4. 强化文件存储的位置跟踪

确保file_storage配置正确,添加超时参数保证读取位置及时持久化:

extensions:
  file_storage:
    directory: /var/addon/splunk/otel_pos
    timeout: 30s # 确保位置信息及时写入磁盘

内容的提问来源于stack exchange,提问作者Paul V

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 03:58:11