You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Stormcrawler结合Elasticsearch示例爬取时Kibana无FETCHED状态页面

Troubleshooting: FETCHED Pages Not Showing in Kibana (Stormcrawler + Elasticsearch)

Let’s work through the most likely fixes for this issue—since your logs confirm pages are being successfully fetched, the problem almost certainly lies in how data flows between Stormcrawler, Elasticsearch, and Kibana.

1. First, Confirm Data Exists in Elasticsearch

Kibana only visualizes data that’s already stored in Elasticsearch, so let’s rule out a write issue first:

  • Run this curl command (replace placeholders with your ES host/port and status index name; the default index for Stormcrawler is usually crawler-status):
    curl -X GET "http://<your-es-host>:9200/<status-index>/_search?q=status:FETCHED"
    
  • If you get zero hits, the problem is with Stormcrawler writing to Elasticsearch. If you do see results, skip straight to step 3.

2. Fix Stormcrawler’s Elasticsearch Status Configuration

If Elasticsearch has no FETCHED entries, dive into your crawler’s crawler-conf.yaml file:

  • Ensure the status updater bolt is enabled (this is what writes crawl statuses to ES):
    status.updater.class: "com.digitalpebble.stormcrawler.elasticsearch.persistence.StatusUpdaterBolt"
    
  • Double-check Elasticsearch connection details are accurate:
    es.status.addresses: ["<your-es-host>"]
    es.status.port: 9200
    es.status.index.name: "crawler-status" # Match this to the index you queried with curl
    
  • Check the Storm UI for errors: Navigate to your Storm cluster’s UI, find your crawler topology, and look at the StatusUpdaterBolt metrics. If you see failed tuples or zero throughput, you might have a connection/permission issue (e.g., your ES cluster requires auth but you haven’t added es.status.username/es.status.password to your config).

3. Fix Kibana Index Pattern Setup

If Elasticsearch has the FETCHED data but Kibana doesn’t display it:

  • Go to Kibana > Stack Management > Index Patterns
  • Make sure you have an index pattern that matches your status index (e.g., crawler-status* if you’re using daily rolling indices)
  • When creating/editing the pattern, select the correct time field (Stormcrawler uses nextFetchDate or fetchTime by default—pick one that’s populated in your ES documents)
  • Refresh the Kibana Discover page and adjust the time range to match when you ran the crawl (it’s easy to accidentally filter out recent data with a narrow time window).

4. Check Elasticsearch’s Refresh Interval

Elasticsearch has a default refresh interval of 1 second, but if this was increased, data might take longer to show up in searches. You can manually refresh the index to test:

curl -X POST "http://<your-es-host>:9200/<status-index>/_refresh"

Then check Kibana again to see if the data appears.


内容的提问来源于stack exchange,提问作者Jonas Pohlmann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:06:42