You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Logstash配置后仍无法重新写入已删除的Elasticsearch日志求助

问题描述

我使用Logstash读取日志文件并发送至Elasticsearch,流式模式运行正常,每日创建不同索引并实时写入日志。昨天下午3点误删索引,索引自动恢复后继续写入,但丢失了当日0点至3点的日志。

为补全全部日志,我删除了sincedb文件,并在Logstash配置中添加ignore_older => 0,再次删除索引后,Logstash仍仅进行流式处理,忽略旧数据。

我的Logstash当前配置如下:

input {
      file {
        path => ["/someDirectory/Logs/20221220-00001.log"]
        start_position => "beginning"
        tags => ["prod"]
        ignore_older => 0
        sincedb_path => "/dev/null"
        type => "cowrie"
      }        
}

filter {    
        grok {
           match => ["path", "/var/www/cap/cap-server/Logs/%{GREEDYDATA:index_name}" ]
        }
}

output {    
       elasticsearch {
          hosts => "IP:9200"
          user => "elastic"
          password => "xxxxxxxx"
          index => "logstash-log-%{index_name}"
       }
}

Elasticsearch配置如下:

# Lock the memory on startup:
#
#bootstrap.memory_lock: true
#
# Make sure that the heap size is set to about half the memory available
# on the system and that the owner of the process is allowed to use this
# limit.
#
# Elasticsearch performs poorly when the system is swapping the memory.
#
# ---------------------------------- Network -----------------------------------
#
# By default Elasticsearch is only accessible on localhost. Set a different
# address here to expose this node on the network:
#
network.host: 0.0.0.0
#
# By default Elasticsearch listens for HTTP traffic on the first free port it
# finds starting at 9200. Set a specific HTTP port here:
#
http.port: 9200
#
# For more information, consult the network module documentation.
#
# --------------------------------- Discovery ----------------------------------
#
# Pass an initial list of hosts to perform discovery when this node is started:
# The default list of hosts is ["127.0.0.1", "[::1]"]
#
#discovery.seed_hosts: ["host1", "host2"]
#
# Bootstrap the cluster using an initial set of master-eligible nodes:
#
#cluster.initial_master_nodes: ["node-1", "node-2"]
#
# For more information, consult the discovery and cluster formation module documentation.
#
# ---------------------------------- Various -----------------------------------
#
# Require explicit names when deleting indices:
#
discovery.type: single-node
xpack.security.enabled: true
xpack.security.transport.ssl.enabled: true
#action.destructive_requires_name: true

注:所有配置修改后,logstash和elasticsearch均已重启。

解决方案

1. 修正路径匹配不一致问题

Logstash input指定的日志路径为/someDirectory/Logs/20221220-00001.log,但filter中grok匹配的路径是/var/www/cap/cap-server/Logs/%{GREEDYDATA:index_name},两者不匹配会导致index_name字段提取失败,直接影响日志写入目标索引。

将grok规则修改为与input路径一致:

filter {    
        grok {
           match => ["path", "/someDirectory/Logs/%{GREEDYDATA:index_name}" ]
        }
}

2. 强制清除Logstash文件读取缓存

即使配置了sincedb_path => "/dev/null",仍可能存在残留缓存导致Logstash跳过旧文件:

  • 停止Logstash服务
  • 检查并删除默认路径下的sincedb文件(通常为$HOME/.sincedb*)
  • 确认目标日志文件的修改时间未设置为未来时间(Logstash会延迟处理修改时间晚于当前时间的文件)
  • 重新启动Logstash服务

3. 验证文件读取权限

确保Logstash运行用户(通常为logstash)拥有目标日志文件的读取权限,执行以下命令检查:

ls -l /someDirectory/Logs/20221220-00001.log

若权限不足,使用chmod或chown调整权限:

chown logstash:logstash /someDirectory/Logs/20221220-00001.log

4. 临时切换输出排查问题

将output临时改为stdout,确认Logstash是否真的读取到旧日志:

output {    
       stdout { codec => rubydebug }
       # 暂时注释Elasticsearch输出
       # elasticsearch {
       #    hosts => "IP:9200"
       #    user => "elastic"
       #    password => "xxxxxxxx"
       #    index => "logstash-log-%{index_name}"
       # }
}

启动Logstash后查看控制台输出:

  • 若能看到旧日志,说明问题出在Elasticsearch写入环节
  • 若看不到旧日志,回到input配置继续排查文件读取逻辑

5. 确认Elasticsearch索引状态

删除索引后,验证新索引是否正常创建并可写入:

curl -u elastic:xxxxxxxx http://IP:9200/logstash-log-20221220-00001/_stats

若索引不存在,说明index_name字段提取失败,回到grok配置修正路径匹配规则。

内容的提问来源于stack exchange,提问作者Tarlan Mammadzada

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 21:40:21