Filebeat日志处理延迟、数据丢失问题排查与解决请求
问题描述
3台主机各运行9个服务,单日志文件每分钟生成30000至72000条事件,服务器配置为16核CPU、62GB内存。当前存在两个核心问题:
- 日志处理延迟约20分钟;
- 每小时日志轮转重命名后,Filebeat不再读取原文件,丢失最后10-20分钟的日志数据。
对比同应用内9台仅配置3-5个日志文件的主机,日志处理高效无异常。已执行GET _cat/thread_pool/bulk?v命令,未发现拒绝请求,bulk线程池运行正常。
环境信息
A] Filebeat/ELK版本
8.3.1(12节点铂金许可证)
B] 每分钟事件数
Service1 - 60k events/min Service2 - 15k events/min Service3 - 500 events/min Service4 - 17k events/min Service5 - 15.5k events/min Service6 - 40k events/min Service7 - 100 events/min Service8 - 78k events/min Service9 - 160k events/min
C] filebeat.yml配置
filebeat.config.modules: enabled: true path: ${path.config}/modules.d/*.yml reload.enabled: true reload.period: 10s queue.mem: events: 20000 #events: 20000 (tried 8192 as well) flush.min_events: 2048 filebeat.registry.flush: 5s output.logstash: hosts: [ "10.40.7.40:5045", "10.40.7.41:5044", "10.40.7.42:5045", "10.40.7.40:5044", "10.40.7.41:5045", "10.40.7.42:5044" ] loadbalance: true workers: 12 #workers: 8 ( tried 8 as well) bulk_max_size: 20000 #bulk_max_size: 15000 (tried this as well) flush_interval: 1s #flush_interval: 500ms (tried this as well)
D] modules.d/app.yml配置
注:每小时日志文件会重命名为"app1-2024-10-16-19-1.log",其中"19"代表上一小时。
- module: app-1 microservice: enabled: true var.paths: - "/path/to/app1.log" input: scan_frequency: 3s close_renamed: false close_inactive: 30m ignore_older: 24h clean_inactive: 48h close_timeout: 2h
E] Logstash主机配置
3台,12核CPU、23GB内存,已设置pipeline.batch_size为2048
F] 应用日志示例
2024-06-13 17:50:20:815 [app4] [app-services] [INFO ] [https-jsse-nio-9443-exec-38] [c7dd483c-2461-4ee5-8d67-bd5076129460] [858f223f1ffe4b8c] [858f223f1ffe4b8c] [] [ResponseLogFilter:100] - TraceId= [c7dd483c-d5076129460] Timestamp= [1718281219341] ClientId= [abcd] AuthMethod= [JOSE] RequestMethod= [POST] ContentType= [application/jose] AcceptType= [application/jose] RequestUri= [/abcd/efg/create] RequestIp= [0.0.0.0] ResponseBody= [{"error_type":"api_validation_error","error_code":"T3","message":"testing test","status":422}] KeyId= [bdbhsbsa-bsbbsbwh8QaCek] Algorithm= [null] ServerAuthorization= [null] Status= [422] ResponseDate= [2024-06-13T17:50:20+0530] ErrorType= [api_validation_error] ErrorCode= [T3] ErrorMessage= [Testing Testing] Latency= [27]
G] Filebeat日志
{"log.level":"info","@timestamp":"2024-10-16T20:23:08.455+0530","log.logger":"monitoring","log.origin":{"file.name":"log/log.go","file.line":185},"message":"Non-zero metrics in the last 30s","service.name":"filebeat","monitoring":{"metrics":{"beat":{"cgroup":{"cpuacct":{"total":{"ns":299461410874}},"memory":{"mem":{"usage":{"bytes":13615104}}}},"cpu":{"system":{"ticks":1140,"time":{"ms":360}},"total":{"ticks":20510,"time":{"ms":7210},"value":0},"user":{"ticks":19370,"time":{"ms":6850}}},"info":{"ephemeral_id":"9dadf2d8-7ad2-43c5-9726-5e33fbb455c8","uptime":{"ms":90114},"version":"8.3.1"},"memstats":{"gc_next":215038376,"memory_alloc":157361792,"memory_sys":4194304,"memory_total":3315695936,"rss":335130624},"runtime":{"goroutines":116}},"filebeat":{"events":{"active":-4,"added":110716,"done":110720},"harvester":{"open_files":0,"running":0}},"libbeat":{"config":{"module":{"running":9},"scans":3},"output":{"events":{"acked":102528,"active":12288,"batches":55,"total":110720},"read":{"bytes":336},"write":{"bytes":22499231}},"pipeline":{"clients":9,"events":{"active":20025,"published":110720,"total":110716},"queue":{"acked":110720}}},"registrar":{"states":{"current":0}},"system":{"load":{"1":13.72,"15":15.51,"5":14.25,"norm":{"1":0.8575,"15":0.9694,"5":0.8906}}}},"ecs.version":"1.6.0"}
解决方案
一、解决日志处理延迟问题
- 优化Filebeat内存队列配置
当前queue.mem.events设置为20000,单台主机每分钟总事件数约386k,每秒约6435条。20000的队列容量仅能缓存3秒事件,易因输出阻塞导致队列满。建议调整:
queue.mem: events: 100000 flush.min_events: 4096 flush.timeout: 5s
增大队列容量同时平衡批量发送效率与延迟。
- 调整Filebeat输出参数
- 将
workers从12调高到16,充分利用16核CPU资源; - 把
bulk_max_size从20000改为8192,配合flush_interval=2s,让批量发送更均匀,降低Logstash瞬时压力; - 开启
compression_level:3,减少网络传输数据量:
output.logstash: hosts: [ "10.40.7.40:5045", "10.40.7.41:5044", "10.40.7.42:5045", "10.40.7.40:5044", "10.40.7.41:5045", "10.40.7.42:5044" ] loadbalance: true workers: 16 bulk_max_size: 8192 flush_interval: 2s compression_level: 3
- 优化Logstash处理能力
- 设置
pipeline.workers:10(12核CPU预留部分资源给系统); - 添加
pipeline.batch.delay:50ms,让Logstash攒够批量再处理; - 调整JVM配置为
-Xms16g -Xmx16g,充分利用23GB内存。
- 降低注册表刷新频率
当前filebeat.registry.flush:5s过于频繁,增加磁盘IO开销,建议改为30s:
filebeat.registry.flush: 30s
二、解决日志轮转后丢失数据问题
从Filebeat日志看harvester.open_files:0和harvester.running:0,说明轮转后采集器被关闭。问题出在close_inactive:30m,原文件停止写入后30分钟采集器关闭,剩余内容未读完。
调整方案:
- 延长
close_inactive时长到1小时,确保轮转后的文件有足够时间被读取:
input: scan_frequency: 3s close_renamed: false close_inactive: 1h ignore_older: 24h clean_inactive: 48h close_timeout: 2h
- 添加
close_removed:false,确保文件重命名后,Filebeat仍保持句柄继续读取:
input: scan_frequency: 3s close_renamed: false close_removed: false close_inactive: 1h ignore_older: 24h clean_inactive: 48h close_timeout: 2h
- 扩展日志路径匹配,让Filebeat识别轮转后的文件:
var.paths: - "/path/to/app1.log" - "/path/to/app1-*.log"
验证步骤
- 应用配置后重启Filebeat;
- 监控Filebeat日志中
filebeat.harvester.running是否保持为9,libbeat.output.events.active是否稳定; - 观察日志处理延迟是否降低,轮转后是否再出现数据丢失;
- 检查Logstash的
pipeline.events.duration_in_millis指标,确认处理效率提升。
内容的提问来源于stack exchange,提问作者Akshay Kulkarni

