关于JSoupParserBolt、Storm Crawler XPath提取及配置的技术咨询
Hey there! Let's break down your questions about Storm Crawler's JSoupParserBolt and configuration step by step:
Absolutely! JSoupParserBolt does come with built-in support for parsefilters.json and its related classes. This file is Storm Crawler's default configuration for parsing filters, which defines a set of rules to process HTML content—like extracting specific elements, cleaning up redundant content, etc. The associated filter classes (such as JSoupFilter and others that implement the ParseFilter interface) are already integrated into Storm Crawler's core module. JSoupParserBolt will automatically load and apply these filters without requiring you to manually implement the basic associated classes.
Filter classes for XPath extraction
You don't need to rewrite filter classes to get specific XPath-based data! Storm Crawler's existing filter system fully supports custom XPath rules. All you need to do is add a custom filter entry in parsefilters.json, specifying the JSoupXPathFilter (or the appropriate XPath-supporting filter class) along with your target XPath expression. JSoupParserBolt will automatically use this filter to extract the data you need. Here's a quick example of what that config might look like:
{ "class": "com.digitalpebble.stormcrawler.parse.filter.JSoupXPathFilter", "name": "CustomContentExtractor", "params": { "xpath": "//article[@id='main-content']/p/text()", "key": "main_article_text" } }
This will extract the text from <p> tags inside the <article id='main-content'> element and store it in the metadata field main_article_text.
Understanding crawler_conf.yaml config items
indexer.md.filter: This config controls which metadata fields are included or excluded during the indexing process. It uses regular expressions with prefixes to define rules:- Use
+to include a field (e.g.,+main_article_textensures this field is sent to the index) - Use
-to exclude a field (e.g.,-.*cookie.*excludes all metadata fields containing "cookie")
If you don't set this config, all metadata fields will be included by default.
- Use
indexer.md.mapping: This config lets you map metadata field names to the field names used in your index. For example, if you want themain_article_textmetadata field to show up ascontentin your index, you'd set this tomain_article_text:content. You can define multiple mappings separated by commas, liketitle:document_title,main_article_text:content.
内容的提问来源于stack exchange,提问作者Abhishek Kumar

