如何用Logstash提取JSON字段并覆盖索引?附示例配置与数据
Extract JSON Fields & Overwrite Elasticsearch Index with Logstash
Got it, let's break down exactly how to handle your JSON array, extract its fields, and set up index overwriting in Logstash. Here's a complete, tested configuration along with explanations for each key part:
Full Logstash Configuration
input { http_poller { urls => { myurl => { method => "GET" url => "http://localhost:8080/helloworld" } } codec => "json" schedule => { cron => "* * * * *" } # Runs every minute, adjust as needed tags => ["helloworld-api"] } } filter { # Split the JSON array into individual events (one per object in the array) split { field => "message" target => "event_data" } # Extract fields from the split object into the top-level of the event mutate { add_field => { "[@metadata][document_id]" => "%{[event_data][key]}" "field1" => "%{[event_data][field1]}" "field2" => "%{[event_data][field2]}" } # Optional: Remove the original nested field if you don't need it remove_field => ["message", "event_data"] } } output { elasticsearch { hosts => ["http://localhost:9200"] index => "helloworld_index" # Your target index name document_id => "%{[@metadata][document_id]}" # Use the unique "key" field as document ID action => "index" # This will overwrite existing documents with the same ID # Optional: Add authentication if your ES instance requires it # user => "elastic" # password => "your_password" } # Optional: Print events to console for debugging stdout { codec => rubydebug } }
Key Explanations
Let's walk through the most important parts that make this work:
1. Input: Fetching the JSON Array
- The
http_pollerinput hits your endpoint and uses thejsoncodec to parse the response automatically. - The
scheduleuses cron syntax (adjust the value to match your desired polling frequency).
2. Filter: Splitting & Extracting Fields
splitfilter: Your API returns a JSON array, so we split it into separate events—one for each{ "key": "...", ... }object. This ensures each item is processed individually.mutatefilter: We extract thekey,field1, andfield2values into the top-level of the event. We also store thekeyin[@metadata][document_id](metadata fields don't get indexed, which keeps your ES clean).
3. Output: Overwriting the Index
- The critical part for overwriting is setting
document_id => "%{[@metadata][document_id]}". Since each object has a uniquekey, this ensures that when a new event with the samekeycomes in, Elasticsearch will replace the existing document with the new data. - Using
action => "index"is the default, but it explicitly tells ES to create or replace the document (which is exactly what we want for overwriting).
Testing Tips
- Enable the
stdoutoutput withrubydebugcodec to verify that fields are being extracted correctly before sending to Elasticsearch. - Check your Elasticsearch index to confirm that documents with the same
keyare being updated instead of duplicated.
内容的提问来源于stack exchange,提问作者Sneijky
相关产品推荐
相关产品推荐

