You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Apache Nutch 1.14中获取HTML的rawContent并编写index-html插件

获取Apache Nutch 1.14中HTML原始内容的方案

我刚好对Nutch 1.x的插件开发比较熟悉,针对你提到的“1.14版本无法像2.3.1那样直接通过WebPage获取HTML原始内容”的问题,给你梳理几个可行的实现思路:

核心差异说明

Nutch 1.x和2.x的存储模型完全不同:2.x依赖外部存储引擎(比如HBase),WebPage对象会持久化抓取到的原始内容;而1.x是基于HDFS的segment文件存储,原始HTML内容保存在segment的fetcher目录下,解析后的文本和元数据才会进入后续的indexing流程。所以要拿到原始内容,得从segment文件里读取。

方法一:在Indexing插件中直接读取Segment的FetcherOutput

这是最直接的方案,在你的自定义index-html插件的filter方法里,通过SegmentReader读取当前处理的segment文件,提取对应URL的原始内容。

代码示例

import org.apache.nutch.indexer.IndexingFilter;
import org.apache.nutch.indexer.NutchDocument;
import org.apache.nutch.parse.Parse;
import org.apache.nutch.segment.SegmentReader;
import org.apache.nutch.storage.WebPage;
import org.apache.nutch.util.Utf8;
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.Text;
import org.apache.nutch.crawl.CrawlDatum;

import java.io.IOException;
import java.nio.charset.StandardCharsets;

public class CustomIndexHtmlFilter implements IndexingFilter {

    private Configuration conf;
    private static final org.slf4j.Logger LOG = org.slf4j.LoggerFactory.getLogger(CustomIndexHtmlFilter.class);

    @Override
    public NutchDocument filter(NutchDocument doc, Parse parse, Text url, WebPage page, CrawlDatum datum) throws IOException {
        // 获取当前处理的segment路径(从MapReduce任务上下文获取)
        Path segmentPath = (Path) conf.get("mapreduce.input.fileinputformat.inputdir");
        if (segmentPath == null) {
            LOG.warn("Segment path not found, skip getting raw HTML for URL: {}", url);
            return doc;
        }

        try (SegmentReader reader = new SegmentReader(conf, segmentPath)) {
            // 通过URL获取FetcherOutput对象
            FetcherOutput output = reader.getFetcherOutput(new Utf8(url.toString()));
            if (output != null) {
                byte[] rawHtmlBytes = output.getContent();
                // 从ParseData中获取页面编码,避免乱码
                String encoding = parse.getData().getContentEncoding();
                if (encoding == null || encoding.isEmpty()) {
                    encoding = StandardCharsets.UTF_8.name();
                }
                String rawHtml = new String(rawHtmlBytes, encoding);
                // 将原始HTML添加到索引文档中
                doc.add("raw_html", rawHtml);
            } else {
                LOG.warn("No FetcherOutput found for URL: {}", url);
            }
        } catch (IOException e) {
            LOG.error("Failed to fetch raw HTML for URL: {}", url, e);
        }
        return doc;
    }

    @Override
    public Configuration getConf() {
        return conf;
    }

    @Override
    public void setConf(Configuration conf) {
        this.conf = conf;
    }
}

注意事项

  • 字符编码:一定要用页面实际的编码来转换字节数组,ParseData里的getContentEncoding()可以拿到抓取时识别的编码,避免硬编码UTF-8导致乱码。
  • 资源关闭:使用try-with-resources自动关闭SegmentReader,避免资源泄漏。
  • 分布式环境适配:mapreduce.input.fileinputformat.inputdir会自动指向当前Map任务处理的segment分片路径,在分布式集群中也能正常工作。

方法二:在Parse阶段缓存原始内容

如果不想在indexing阶段读取segment文件,也可以自定义一个Parse插件,在解析之前把原始HTML内容存入ParseData的元数据中,后续indexing插件直接从ParseData中读取。

步骤

  1. 编写自定义ParseFilter,在filter方法中获取原始内容(需要从FetcherOutput读取,或者在ParseContext中拿到原始字节)。
  2. 将原始HTML存入ParseData.getMeta()中,比如parseData.getMeta().put("raw_html", rawHtmlString)。
  3. 在你的index-html插件中,直接通过parse.getData().getMeta().get("raw_html")获取内容。

这个方法的优势是indexing阶段无需读取segment,但需要额外开发Parse插件,流程稍复杂。

总结

如果只是在indexing阶段需要原始HTML,推荐用方法一,直接读取segment的FetcherOutput是最简洁高效的方案,不需要修改太多现有流程。

内容的提问来源于stack exchange,提问作者Kp88

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:19:49