Apache Nutch 2.3.1如何优先抓取种子域名并按比例分配URL?
Is This Feasible?
Absolutely! Your hunch that the Generator component is the right place to adjust this is 100% correct. Nutch's Generator is responsible for picking URLs from the CrawlDb for each crawl iteration, so we can modify its logic to enforce your desired ratio perfectly.
Step-by-Step Implementation
1. First: Define Your Seed Domains Clearly
First, make sure Nutch can reliably identify URLs from your seed domains.
- Add your seed domains to a custom config in
nutch-site.xmlso our custom logic can reference them easily:<property> <name>crawl.seed.domains</name> <value>your-seed-domain.com,another-seed.org</value> <description>Comma-separated list of domains from your seed file</description> </property> - Double-check your
regex-urlfilter.txtto ensure it doesn't accidentally block or allow URLs in a way that breaks this classification.
2. Build a Custom URL Selector (The Core Fix)
Nutch uses a URLSelector to pick URLs in the Generator. The default selector just sorts by score, so we need to build a custom one that splits URLs into seed vs external groups and picks them in your desired ratio.
Here's a simplified implementation you can adapt:
package org.apache.nutch.crawl; import org.apache.hadoop.conf.Configuration; import org.apache.nutch.util.URLUtil; import java.util.*; import java.util.Map.Entry; public class RatioBasedURLSelector implements URLSelector { private Configuration conf; private float seedUrlRatio = 0.6f; // Default to 60% seed URLs private Set<String> seedDomains; @Override public void setConf(Configuration conf) { this.conf = conf; // Pull ratio from config so you can tweak it without recompiling seedUrlRatio = conf.getFloat("crawl.seed.url.ratio", 0.6f); // Load seed domains from our custom config String seedDomainsStr = conf.get("crawl.seed.domains"); if (seedDomainsStr != null) { seedDomains = new HashSet<>(Arrays.asList(seedDomainsStr.split(","))); } else { seedDomains = new HashSet<>(); } } @Override public List<Entry<String, CrawlDatum>> select(Map<String, CrawlDatum> crawlDb, int numUrlsToPick) { // Split URLs into seed domain and external groups List<Entry<String, CrawlDatum>> seedUrls = new ArrayList<>(); List<Entry<String, CrawlDatum>> externalUrls = new ArrayList<>(); for (Entry<String, CrawlDatum> entry : crawlDb.entrySet()) { String url = entry.getKey(); try { String domain = URLUtil.getDomainName(url); if (seedDomains.contains(domain.trim())) { seedUrls.add(entry); } else { externalUrls.add(entry); } } catch (Exception e) { // Skip malformed URLs continue; } } // Calculate how many URLs to pick from each group int seedCount = Math.round(numUrlsToPick * seedUrlRatio); int externalCount = numUrlsToPick - seedCount; // Don't pick more URLs than exist in each group seedCount = Math.min(seedCount, seedUrls.size()); externalCount = Math.min(externalCount, externalUrls.size()); // Keep the default behavior of picking higher-score URLs first seedUrls.sort((a, b) -> Float.compare(b.getValue().getScore(), a.getValue().getScore())); externalUrls.sort((a, b) -> Float.compare(b.getValue().getScore(), a.getValue().getScore())); // Combine the lists List<Entry<String, CrawlDatum>> selectedUrls = new ArrayList<>(); selectedUrls.addAll(seedUrls.subList(0, seedCount)); selectedUrls.addAll(externalUrls.subList(0, externalCount)); // Optional: Shuffle the combined list to avoid batch-crawling all seed URLs first Collections.shuffle(selectedUrls); return selectedUrls; } @Override public Configuration getConf() { return conf; } }
3. Configure Nutch to Use Your Custom Selector
Update nutch-site.xml to tell the Generator to use your new selector, and set your desired ratio:
<property> <name>generator.selector.class</name> <value>org.apache.nutch.crawl.RatioBasedURLSelector</value> <description>Custom selector to enforce seed vs external URL crawl ratio</description> </property> <property> <name>crawl.seed.url.ratio</name> <value>0.6</value> <description>Ratio of seed domain URLs to pick (0.6 = 60%), the rest are external</description> </property>
4. Compile and Test
- Compile your custom class and add it to Nutch's classpath (either by packaging it into a JAR and dropping it in
lib/, or rebuilding Nutch from source if you modified the core code). - Run a test crawl iteration, then check the generated fetchlist to verify the ratio. If seed URLs are scarce, the selector will automatically grab all available seed URLs and fill the rest with external ones.
- Tweak the
crawl.seed.url.ratiovalue if you need to adjust the balance later.
Quick Notes
- If your seed domain has fewer URLs than the 60% quota, the selector will take all available seed URLs and use external URLs for the remaining slots.
- Keeping the score-sorting logic ensures you're still prioritizing higher-quality URLs within each group, not just random ones.
- The shuffle step is optional but helps avoid crawler bias where it crawls all seed URLs first then switches to externals, which might trigger anti-bot measures on some sites.
内容的提问来源于stack exchange,提问作者Hafiz Muhammad Shafiq

