如何使用Logstash将Elasticsearch索引导出为CSV至Google Cloud Storage
解决Elasticsearch数据导出CSV到Google Cloud Storage的问题
你的现有配置存在结构错误:把Logstash的filter块写到了Elasticsearch input的query DSL语句里,这会导致配置解析失败,无法正确完成字段重命名和CSV导出。以下是修正后的完整配置:
input { elasticsearch { hosts => "localhost:9200" # 按天匹配索引,替换成你的按天索引命名规则,比如test-2024.05.20 index => "test-%{+YYYY.MM.dd}" query => '{ "_source": ["field1","field2"], "query": { "match_all": {} } }' # 如果需要定时同步每天数据,可以添加schedule参数,比如每天凌晨1点执行 # schedule => "0 1 * * *" } } filter { mutate { rename => { "field1" => "test1" "field2" => "test2" } } } output { google_cloud_storage { codec => csv { include_headers => true columns => ["test1", "test2"] } bucket => "bucketName" json_key_file => "creds.json" temp_directory => "/tmp" log_file_prefix => "logstash_gcs" max_file_size_kbytes => 1024 date_pattern => "%Y-%m-%dT%H:00" flush_interval_secs => 600 gzip => false uploader_interval_secs => 600 include_uuid => true include_hostname => true } }
关键修正说明:
- 索引匹配:将
index设置为test-%{+YYYY.MM.dd},让Logstash自动匹配当天的按天索引(如果是导出历史某天数据,直接写具体索引名如test-2024.05.20即可)。 - 字段重命名:把字段重命名逻辑移到独立的
filter块中,这是Logstash处理字段的正确阶段,不能嵌套在ES的query里。 - CSV编码:
csv codec的columns参数指定了导出的字段顺序,include_headers: true会在CSV文件开头生成表头行。 - GCS配置:确保
json_key_file的路径正确指向你的GCS服务账号密钥文件,bucket填写实际的存储桶名称。
如果需要定时每天自动导出前一天的索引,可以调整index为test-%{+YYYY.MM.dd-1d},并配合schedule参数设置执行时间。
内容的提问来源于stack exchange,提问作者Amulya M
相关产品推荐
相关产品推荐

