使用Nutch爬取仅获取单版本URL的问题解决咨询
Hey there, the problem you're hitting is due to Nutch's default URL deduplication and normalization rules—it's treating URLs like example.com/docs?v=1 and example.com/docs?v=2 as duplicates, even though their content differs. Here's how to fix this step by step:
1. Adjust URL Normalization to Keep the v Parameter
Nutch's default URL normalizer might strip or ignore query parameters that it doesn't recognize as meaningful. To make sure the v parameter is preserved:
- Navigate to your Nutch config directory and open
regex-normalize.xml. - Add or update a rule to explicitly retain the
vparameter:
This ensures URLs with different<regex> <pattern>^(.*)(\?.*v=\d+)(.*)$</pattern> <replace>$1$2$3</replace> </regex>vvalues won't get normalized into the same path.
2. Switch to Full URL-Based Deduplication
By default, Nutch might use a path-based deduplication policy. To make it consider the full URL (including query params) when checking for duplicates:
- Open
nutch-site.xmland set thedb.duplicate.policy.classproperty to use the URL-based policy:
This class treats two URLs as distinct if their full string (including query parameters) is different, so your two versioned pages will be seen as separate entries.<property> <name>db.duplicate.policy.class</name> <value>org.apache.nutch.crawl.URLDuplicatePolicy</value> </property>
3. Update URL Filter Rules to Allow Versioned URLs
Make sure your URL filter isn't blocking URLs with the v parameter:
- Open
regex-urlfilter.txtin the config directory. - Add a rule to explicitly allow URLs from your domain with the
vparameter:
Or, if you want to allow all URLs from your domain (including those with any query params), use:+^https?://example.com/.*\?v=\d+$
Just ensure there are no conflicting+^https?://example.com/.*-rules that might exclude these URLs.
4. Clean Up Old Crawl Data and Re-run
Before re-running your crawl, delete the existing crawl directory to avoid conflicts with old deduplication data. Then execute your original command again:
bin/crawl -s urls crawl 3
Now Nutch should crawl both v=1 and v=2 versions of your pages as separate, unique URLs.
If you need more control over deduplication later, you can even implement a custom DuplicatePolicy class, but the steps above should resolve your immediate issue.
内容的提问来源于stack exchange,提问作者Neenu Chandran

