Nutch-1.15中能否单独调用extractor插件处理已解析数据?
Great question! I’ve worked with Nutch’s extractor plugin extensively, and yes—you absolutely can run it standalone against already parsed data without restarting your entire crawl pipeline. Here’s how to do it step by step:
1. Locate Your Parsed Segment Data
Nutch stores parsed content in its segment directories. Each segment contains a parse_data subdirectory with parsed records (in SequenceFile format). This is the data we’ll target for standalone extraction.
2. Use Nutch’s ParseSegment Tool with Extractor Configuration
Nutch’s built-in ParseSegment tool can be adapted to run just the extractor logic on existing parsed data. Here’s a command to execute this:
bin/nutch parsesegment -Dparser.extractor.config.file=conf/custom-extractors.xml -Dplugin.includes=extractor -extract /path/to/your/target/segment
Let’s break down the flags:
-Dparser.extractor.config.file: Points to your updatedcustom-extractors.xml(use the full path if it’s not in the defaultconf/directory).-Dplugin.includes=extractor: Ensures the extractor plugin is loaded and used.-extract: Triggers the extraction step on the existing parsed data in the segment.
3. Validate the Extracted Results
After running the command, you can verify the output using Nutch’s ReadSegment tool to inspect the extracted fields:
bin/nutch readsegment -dump /path/to/your/target/segment /path/to/output/dump -extract
This will dump the extracted data to the specified output directory, so you can confirm your updated rules are working as expected.
4. Batch Process Multiple Segments
If you have multiple segments to reprocess, create a simple shell script to loop through all segment directories and run the ParseSegment command on each. For example:
#!/bin/bash SEGMENT_DIR="/path/to/all/segments" for seg in $SEGMENT_DIR/*; do if [ -d "$seg" ]; then echo "Processing segment: $seg" bin/nutch parsesegment -Dparser.extractor.config.file=conf/custom-extractors.xml -Dplugin.includes=extractor -extract "$seg" fi done
This script will automate reprocessing all segments with your updated extractor rules in one go.
Important Notes
- Ensure your
custom-extractors.xmlis correctly formatted—even minor typos can break the extraction logic, so double-check before running the command. - The extractor plugin depends on valid parsed data (like HTML DOM or plain text) in the segment. As long as the
parse_datadirectory exists and contains valid records, standalone processing will work. - This approach avoids restarting the entire crawl flow—you only reprocess the segments that need the updated extraction rules.
内容的提问来源于stack exchange,提问作者Sudha Sompura

