You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Apache Nutch 2.3.1如何优先抓取种子域名并按比例分配URL?

Controlling Crawl Ratio: 60% Seed Domain URLs, 40% External URLs

Is This Feasible?

Absolutely! Your hunch that the Generator component is the right place to adjust this is 100% correct. Nutch's Generator is responsible for picking URLs from the CrawlDb for each crawl iteration, so we can modify its logic to enforce your desired ratio perfectly.

Step-by-Step Implementation

1. First: Define Your Seed Domains Clearly

First, make sure Nutch can reliably identify URLs from your seed domains.

  • Add your seed domains to a custom config in nutch-site.xml so our custom logic can reference them easily:
    <property>
      <name>crawl.seed.domains</name>
      <value>your-seed-domain.com,another-seed.org</value>
      <description>Comma-separated list of domains from your seed file</description>
    </property>
    
  • Double-check your regex-urlfilter.txt to ensure it doesn't accidentally block or allow URLs in a way that breaks this classification.

2. Build a Custom URL Selector (The Core Fix)

Nutch uses a URLSelector to pick URLs in the Generator. The default selector just sorts by score, so we need to build a custom one that splits URLs into seed vs external groups and picks them in your desired ratio.

Here's a simplified implementation you can adapt:

package org.apache.nutch.crawl;

import org.apache.hadoop.conf.Configuration;
import org.apache.nutch.util.URLUtil;
import java.util.*;
import java.util.Map.Entry;

public class RatioBasedURLSelector implements URLSelector {
    private Configuration conf;
    private float seedUrlRatio = 0.6f; // Default to 60% seed URLs
    private Set<String> seedDomains;

    @Override
    public void setConf(Configuration conf) {
        this.conf = conf;
        // Pull ratio from config so you can tweak it without recompiling
        seedUrlRatio = conf.getFloat("crawl.seed.url.ratio", 0.6f);
        // Load seed domains from our custom config
        String seedDomainsStr = conf.get("crawl.seed.domains");
        if (seedDomainsStr != null) {
            seedDomains = new HashSet<>(Arrays.asList(seedDomainsStr.split(",")));
        } else {
            seedDomains = new HashSet<>();
        }
    }

    @Override
    public List<Entry<String, CrawlDatum>> select(Map<String, CrawlDatum> crawlDb, int numUrlsToPick) {
        // Split URLs into seed domain and external groups
        List<Entry<String, CrawlDatum>> seedUrls = new ArrayList<>();
        List<Entry<String, CrawlDatum>> externalUrls = new ArrayList<>();

        for (Entry<String, CrawlDatum> entry : crawlDb.entrySet()) {
            String url = entry.getKey();
            try {
                String domain = URLUtil.getDomainName(url);
                if (seedDomains.contains(domain.trim())) {
                    seedUrls.add(entry);
                } else {
                    externalUrls.add(entry);
                }
            } catch (Exception e) {
                // Skip malformed URLs
                continue;
            }
        }

        // Calculate how many URLs to pick from each group
        int seedCount = Math.round(numUrlsToPick * seedUrlRatio);
        int externalCount = numUrlsToPick - seedCount;

        // Don't pick more URLs than exist in each group
        seedCount = Math.min(seedCount, seedUrls.size());
        externalCount = Math.min(externalCount, externalUrls.size());

        // Keep the default behavior of picking higher-score URLs first
        seedUrls.sort((a, b) -> Float.compare(b.getValue().getScore(), a.getValue().getScore()));
        externalUrls.sort((a, b) -> Float.compare(b.getValue().getScore(), a.getValue().getScore()));

        // Combine the lists
        List<Entry<String, CrawlDatum>> selectedUrls = new ArrayList<>();
        selectedUrls.addAll(seedUrls.subList(0, seedCount));
        selectedUrls.addAll(externalUrls.subList(0, externalCount));

        // Optional: Shuffle the combined list to avoid batch-crawling all seed URLs first
        Collections.shuffle(selectedUrls);

        return selectedUrls;
    }

    @Override
    public Configuration getConf() {
        return conf;
    }
}

3. Configure Nutch to Use Your Custom Selector

Update nutch-site.xml to tell the Generator to use your new selector, and set your desired ratio:

<property>
  <name>generator.selector.class</name>
  <value>org.apache.nutch.crawl.RatioBasedURLSelector</value>
  <description>Custom selector to enforce seed vs external URL crawl ratio</description>
</property>

<property>
  <name>crawl.seed.url.ratio</name>
  <value>0.6</value>
  <description>Ratio of seed domain URLs to pick (0.6 = 60%), the rest are external</description>
</property>

4. Compile and Test

  • Compile your custom class and add it to Nutch's classpath (either by packaging it into a JAR and dropping it in lib/, or rebuilding Nutch from source if you modified the core code).
  • Run a test crawl iteration, then check the generated fetchlist to verify the ratio. If seed URLs are scarce, the selector will automatically grab all available seed URLs and fill the rest with external ones.
  • Tweak the crawl.seed.url.ratio value if you need to adjust the balance later.

Quick Notes

  • If your seed domain has fewer URLs than the 60% quota, the selector will take all available seed URLs and use external URLs for the remaining slots.
  • Keeping the score-sorting logic ensures you're still prioritizing higher-quality URLs within each group, not just random ones.
  • The shuffle step is optional but helps avoid crawler bias where it crawls all seed URLs first then switches to externals, which might trigger anti-bot measures on some sites.

内容的提问来源于stack exchange,提问作者Hafiz Muhammad Shafiq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:35:23