Logstash配置后仍无法重新写入已删除的Elasticsearch日志求助
我使用Logstash读取日志文件并发送至Elasticsearch,流式模式运行正常,每日创建不同索引并实时写入日志。昨天下午3点误删索引,索引自动恢复后继续写入,但丢失了当日0点至3点的日志。
为补全全部日志,我删除了sincedb文件,并在Logstash配置中添加ignore_older => 0,再次删除索引后,Logstash仍仅进行流式处理,忽略旧数据。
我的Logstash当前配置如下:
input { file { path => ["/someDirectory/Logs/20221220-00001.log"] start_position => "beginning" tags => ["prod"] ignore_older => 0 sincedb_path => "/dev/null" type => "cowrie" } } filter { grok { match => ["path", "/var/www/cap/cap-server/Logs/%{GREEDYDATA:index_name}" ] } } output { elasticsearch { hosts => "IP:9200" user => "elastic" password => "xxxxxxxx" index => "logstash-log-%{index_name}" } }
Elasticsearch配置如下:
# Lock the memory on startup: # #bootstrap.memory_lock: true # # Make sure that the heap size is set to about half the memory available # on the system and that the owner of the process is allowed to use this # limit. # # Elasticsearch performs poorly when the system is swapping the memory. # # ---------------------------------- Network ----------------------------------- # # By default Elasticsearch is only accessible on localhost. Set a different # address here to expose this node on the network: # network.host: 0.0.0.0 # # By default Elasticsearch listens for HTTP traffic on the first free port it # finds starting at 9200. Set a specific HTTP port here: # http.port: 9200 # # For more information, consult the network module documentation. # # --------------------------------- Discovery ---------------------------------- # # Pass an initial list of hosts to perform discovery when this node is started: # The default list of hosts is ["127.0.0.1", "[::1]"] # #discovery.seed_hosts: ["host1", "host2"] # # Bootstrap the cluster using an initial set of master-eligible nodes: # #cluster.initial_master_nodes: ["node-1", "node-2"] # # For more information, consult the discovery and cluster formation module documentation. # # ---------------------------------- Various ----------------------------------- # # Require explicit names when deleting indices: # discovery.type: single-node xpack.security.enabled: true xpack.security.transport.ssl.enabled: true #action.destructive_requires_name: true
注:所有配置修改后,logstash和elasticsearch均已重启。
1. 修正路径匹配不一致问题
Logstash input指定的日志路径为/someDirectory/Logs/20221220-00001.log,但filter中grok匹配的路径是/var/www/cap/cap-server/Logs/%{GREEDYDATA:index_name},两者不匹配会导致index_name字段提取失败,直接影响日志写入目标索引。
将grok规则修改为与input路径一致:
filter { grok { match => ["path", "/someDirectory/Logs/%{GREEDYDATA:index_name}" ] } }
2. 强制清除Logstash文件读取缓存
即使配置了sincedb_path => "/dev/null",仍可能存在残留缓存导致Logstash跳过旧文件:
- 停止Logstash服务
- 检查并删除默认路径下的sincedb文件(通常为
$HOME/.sincedb*) - 确认目标日志文件的修改时间未设置为未来时间(Logstash会延迟处理修改时间晚于当前时间的文件)
- 重新启动Logstash服务
3. 验证文件读取权限
确保Logstash运行用户(通常为logstash)拥有目标日志文件的读取权限,执行以下命令检查:
ls -l /someDirectory/Logs/20221220-00001.log
若权限不足,使用chmod或chown调整权限:
chown logstash:logstash /someDirectory/Logs/20221220-00001.log
4. 临时切换输出排查问题
将output临时改为stdout,确认Logstash是否真的读取到旧日志:
output { stdout { codec => rubydebug } # 暂时注释Elasticsearch输出 # elasticsearch { # hosts => "IP:9200" # user => "elastic" # password => "xxxxxxxx" # index => "logstash-log-%{index_name}" # } }
启动Logstash后查看控制台输出:
- 若能看到旧日志,说明问题出在Elasticsearch写入环节
- 若看不到旧日志,回到input配置继续排查文件读取逻辑
5. 确认Elasticsearch索引状态
删除索引后,验证新索引是否正常创建并可写入:
curl -u elastic:xxxxxxxx http://IP:9200/logstash-log-20221220-00001/_stats
若索引不存在,说明index_name字段提取失败,回到grok配置修正路径匹配规则。
内容的提问来源于stack exchange,提问作者Tarlan Mammadzada

