如何用StormCrawler将网页内容存入Status Index?配置后无数据求指导
Let's walk through exactly what you need to do to get the crawled page content showing up in your Status Index. Here's a step-by-step guide tailored to StormCrawler's codebase:
1. Locate the Core Status Update Class
The component responsible for writing data to the Elasticsearch Status Index is EsStatusUpdaterBolt, found in the org.apache.stormcrawler.elasticsearch package of the stormcrawler-elasticsearch module. This is where you'll make your key code changes.
2. Add Content Field to the ES Document
Open EsStatusUpdaterBolt and find the method that constructs the Elasticsearch document from a CrawlURI (typically named updateDocument). You need to add logic to pull the page content from the CrawlURI and include it in the document:
- Retrieve the raw content using
crawlURI.getContent()(returns abyte[]), then convert it to a UTF-8 string. - Add this string to the document under the "content" field—make sure the field name matches exactly what you defined in your
ES_IndexInit.shmapping.
Example code snippet to insert:
import java.nio.charset.StandardCharsets; // Inside the document-building method if (crawlURI.getContent() != null) { String pageContent = new String(crawlURI.getContent(), StandardCharsets.UTF_8); doc.addField("content", pageContent); }
3. Confirm Content is Passed to the Bolt
Before recompiling, double-check your topology configuration:
- Ensure your
ParserBolt(or content-extracting component) is set up to populate theCrawlURI's content field. - Make sure no intermediate bolts are stripping the content from the
CrawlURIbefore it reachesEsStatusUpdaterBolt.
4. Recompile StormCrawler
With the code changes in place, rebuild the stormcrawler-elasticsearch module (or full project) using Maven:
mvn clean package -DskipTests
This generates updated JAR files that include your custom logic.
5. Update and Restart Your Storm Topology
- Replace the old StormCrawler Elasticsearch JAR in your topology's dependencies with the newly compiled version.
- Kill any running instance of your topology, then submit the updated version to your Storm cluster.
6. Validate the Results
Run your crawl job again, then navigate to Kibana to inspect the Status Index. You should now see the "content" field populated with the crawled page text.
Troubleshooting Tips
- Check Logs: Add debug logging in
EsStatusUpdaterBoltto print the full document being sent to Elasticsearch. This confirms if the "content" field is actually included. - Mapping Match: Verify the field name in your code exactly matches the "content" field defined in your ES mapping (case-sensitive!).
- Large Content Handling: If pages have extremely long content, adjust your ES mapping (e.g., increase
ignore_abovefor the text field) to avoid truncation.
内容的提问来源于stack exchange,提问作者ArtoriasSnow

