You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用StormCrawler将网页内容存入Status Index?配置后无数据求指导

How to Modify StormCrawler to Store Page Content in Elasticsearch Status Index

Let's walk through exactly what you need to do to get the crawled page content showing up in your Status Index. Here's a step-by-step guide tailored to StormCrawler's codebase:

1. Locate the Core Status Update Class

The component responsible for writing data to the Elasticsearch Status Index is EsStatusUpdaterBolt, found in the org.apache.stormcrawler.elasticsearch package of the stormcrawler-elasticsearch module. This is where you'll make your key code changes.

2. Add Content Field to the ES Document

Open EsStatusUpdaterBolt and find the method that constructs the Elasticsearch document from a CrawlURI (typically named updateDocument). You need to add logic to pull the page content from the CrawlURI and include it in the document:

  • Retrieve the raw content using crawlURI.getContent() (returns a byte[]), then convert it to a UTF-8 string.
  • Add this string to the document under the "content" field—make sure the field name matches exactly what you defined in your ES_IndexInit.sh mapping.

Example code snippet to insert:

import java.nio.charset.StandardCharsets;

// Inside the document-building method
if (crawlURI.getContent() != null) {
    String pageContent = new String(crawlURI.getContent(), StandardCharsets.UTF_8);
    doc.addField("content", pageContent);
}

3. Confirm Content is Passed to the Bolt

Before recompiling, double-check your topology configuration:

  • Ensure your ParserBolt (or content-extracting component) is set up to populate the CrawlURI's content field.
  • Make sure no intermediate bolts are stripping the content from the CrawlURI before it reaches EsStatusUpdaterBolt.

4. Recompile StormCrawler

With the code changes in place, rebuild the stormcrawler-elasticsearch module (or full project) using Maven:

mvn clean package -DskipTests

This generates updated JAR files that include your custom logic.

5. Update and Restart Your Storm Topology

  • Replace the old StormCrawler Elasticsearch JAR in your topology's dependencies with the newly compiled version.
  • Kill any running instance of your topology, then submit the updated version to your Storm cluster.

6. Validate the Results

Run your crawl job again, then navigate to Kibana to inspect the Status Index. You should now see the "content" field populated with the crawled page text.

Troubleshooting Tips

  • Check Logs: Add debug logging in EsStatusUpdaterBolt to print the full document being sent to Elasticsearch. This confirms if the "content" field is actually included.
  • Mapping Match: Verify the field name in your code exactly matches the "content" field defined in your ES mapping (case-sensitive!).
  • Large Content Handling: If pages have extremely long content, adjust your ES mapping (e.g., increase ignore_above for the text field) to avoid truncation.

内容的提问来源于stack exchange,提问作者ArtoriasSnow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:33:44