如何将Nutch爬取的网页数据存储至HDFS?已部署Hadoop2.x与Nutch1.x
Hey there! I’ve messed around with Nutch 1.x and Hadoop 2.x a fair bit, so let me walk you through exactly how to get your crawled web data into HDFS. Here’s a step-by-step breakdown:
First, make sure Nutch knows how to talk to your HDFS cluster. Copy the core-site.xml and hdfs-site.xml files from your Hadoop installation’s conf directory into Nutch’s conf folder. This tells Nutch your NameNode address, HDFS port, and other critical cluster details.
Open up Nutch’s nutch-site.xml (in the conf directory) and double-check or add these key settings:
http.agent.name: Set this to a unique identifier for your crawler (e.g.,MyNutchCrawler— this is required to avoid getting blocked by websites).- Ensure the
fs.default.nameproperty matches what’s in your Hadoopcore-site.xml(e.g.,hdfs://your-nn-host:9000). This sets Nutch’s default filesystem to HDFS instead of local disk.
When you kick off a crawl, specify an HDFS path as your crawl directory. For example:
bin/nutch crawl urls -dir hdfs://your-nn-host:9000/my-nutch-crawl -depth 2 -topN 50
This will store all crawl data (segments, crawl DB, link DB, etc.) directly in HDFS under /my-nutch-crawl. The raw crawled page content lives inside the segments subdirectories — each segment contains the pages fetched during a single crawl iteration.
If you want to extract just the page content into plain files in HDFS (like HTML or text), use Nutch’s dump command. For example, to dump content from a specific segment:
bin/nutch dump -segment hdfs://your-nn-host:9000/my-nutch-crawl/segments/202405201430 -output hdfs://your-nn-host:9000/nutch-page-dumps -format html
This will export all pages from that segment into HDFS files in the /nutch-page-dumps directory. You can switch -format html to -format text if you only want the raw text content.
Check that your data is in HDFS using basic Hadoop commands:
# List the crawl directory structure hdfs dfs -ls hdfs://your-nn-host:9000/my-nutch-crawl # Preview a dumped page file hdfs dfs -cat hdfs://your-nn-host:9000/nutch-page-dumps/part-00000 | head -100
Quick Troubleshooting Tips
- If Nutch can’t connect to HDFS, confirm your Hadoop cluster is running (
jpsshould show NameNode, DataNode, etc.), and that the firewall isn’t blocking the HDFS port (usually 9000). - Make sure the user running Nutch has write permissions on the HDFS directory you’re using. You can set this with:
hdfs dfs -mkdir /my-nutch-crawl hdfs dfs -chmod 755 /my-nutch-crawl
内容的提问来源于stack exchange,提问作者Murthy Routhula

