You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Logstash提取JSON字段并覆盖索引?附示例配置与数据

Extract JSON Fields & Overwrite Elasticsearch Index with Logstash

Got it, let's break down exactly how to handle your JSON array, extract its fields, and set up index overwriting in Logstash. Here's a complete, tested configuration along with explanations for each key part:

Full Logstash Configuration

input {
  http_poller {
    urls => {
      myurl => {
        method => "GET"
        url => "http://localhost:8080/helloworld"
      }
    }
    codec => "json"
    schedule => { cron => "* * * * *" } # Runs every minute, adjust as needed
    tags => ["helloworld-api"]
  }
}

filter {
  # Split the JSON array into individual events (one per object in the array)
  split {
    field => "message"
    target => "event_data"
  }

  # Extract fields from the split object into the top-level of the event
  mutate {
    add_field => {
      "[@metadata][document_id]" => "%{[event_data][key]}"
      "field1" => "%{[event_data][field1]}"
      "field2" => "%{[event_data][field2]}"
    }
    # Optional: Remove the original nested field if you don't need it
    remove_field => ["message", "event_data"]
  }
}

output {
  elasticsearch {
    hosts => ["http://localhost:9200"]
    index => "helloworld_index" # Your target index name
    document_id => "%{[@metadata][document_id]}" # Use the unique "key" field as document ID
    action => "index" # This will overwrite existing documents with the same ID
    # Optional: Add authentication if your ES instance requires it
    # user => "elastic"
    # password => "your_password"
  }

  # Optional: Print events to console for debugging
  stdout {
    codec => rubydebug
  }
}

Key Explanations

Let's walk through the most important parts that make this work:

1. Input: Fetching the JSON Array

  • The http_poller input hits your endpoint and uses the json codec to parse the response automatically.
  • The schedule uses cron syntax (adjust the value to match your desired polling frequency).

2. Filter: Splitting & Extracting Fields

  • split filter: Your API returns a JSON array, so we split it into separate events—one for each { "key": "...", ... } object. This ensures each item is processed individually.
  • mutate filter: We extract the key, field1, and field2 values into the top-level of the event. We also store the key in [@metadata][document_id] (metadata fields don't get indexed, which keeps your ES clean).

3. Output: Overwriting the Index

  • The critical part for overwriting is setting document_id => "%{[@metadata][document_id]}". Since each object has a unique key, this ensures that when a new event with the same key comes in, Elasticsearch will replace the existing document with the new data.
  • Using action => "index" is the default, but it explicitly tells ES to create or replace the document (which is exactly what we want for overwriting).

Testing Tips

  • Enable the stdout output with rubydebug codec to verify that fields are being extracted correctly before sending to Elasticsearch.
  • Check your Elasticsearch index to confirm that documents with the same key are being updated instead of duplicated.

内容的提问来源于stack exchange,提问作者Sneijky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:07:44