You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Nutch爬取的网页数据存储至HDFS?已部署Hadoop2.x与Nutch1.x

Hey there! I’ve messed around with Nutch 1.x and Hadoop 2.x a fair bit, so let me walk you through exactly how to get your crawled web data into HDFS. Here’s a step-by-step breakdown:

Step 1: Sync Hadoop Configs to Nutch

First, make sure Nutch knows how to talk to your HDFS cluster. Copy the core-site.xml and hdfs-site.xml files from your Hadoop installation’s conf directory into Nutch’s conf folder. This tells Nutch your NameNode address, HDFS port, and other critical cluster details.

Step 2: Tweak Nutch’s Core Configuration

Open up Nutch’s nutch-site.xml (in the conf directory) and double-check or add these key settings:

  • http.agent.name: Set this to a unique identifier for your crawler (e.g., MyNutchCrawler — this is required to avoid getting blocked by websites).
  • Ensure the fs.default.name property matches what’s in your Hadoop core-site.xml (e.g., hdfs://your-nn-host:9000). This sets Nutch’s default filesystem to HDFS instead of local disk.
Step 3: Run Crawls Directly to HDFS

When you kick off a crawl, specify an HDFS path as your crawl directory. For example:

bin/nutch crawl urls -dir hdfs://your-nn-host:9000/my-nutch-crawl -depth 2 -topN 50

This will store all crawl data (segments, crawl DB, link DB, etc.) directly in HDFS under /my-nutch-crawl. The raw crawled page content lives inside the segments subdirectories — each segment contains the pages fetched during a single crawl iteration.

Step 4: Export Raw Page Content to HDFS (Optional)

If you want to extract just the page content into plain files in HDFS (like HTML or text), use Nutch’s dump command. For example, to dump content from a specific segment:

bin/nutch dump -segment hdfs://your-nn-host:9000/my-nutch-crawl/segments/202405201430 -output hdfs://your-nn-host:9000/nutch-page-dumps -format html

This will export all pages from that segment into HDFS files in the /nutch-page-dumps directory. You can switch -format html to -format text if you only want the raw text content.

Step 5: Verify Everything Worked

Check that your data is in HDFS using basic Hadoop commands:

# List the crawl directory structure
hdfs dfs -ls hdfs://your-nn-host:9000/my-nutch-crawl

# Preview a dumped page file
hdfs dfs -cat hdfs://your-nn-host:9000/nutch-page-dumps/part-00000 | head -100

Quick Troubleshooting Tips

  • If Nutch can’t connect to HDFS, confirm your Hadoop cluster is running (jps should show NameNode, DataNode, etc.), and that the firewall isn’t blocking the HDFS port (usually 9000).
  • Make sure the user running Nutch has write permissions on the HDFS directory you’re using. You can set this with:
    hdfs dfs -mkdir /my-nutch-crawl
    hdfs dfs -chmod 755 /my-nutch-crawl
    

内容的提问来源于stack exchange,提问作者Murthy Routhula

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:17:36