You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

node-crawler与simplecrawler技术咨询:仅爬取内部URL相关疑问

Node.js Crawler Tools: node-crawler & simplecrawler for Internal URL Collection

Hey there! Since you're new to Node.js web crawlers and only need to collect internal URLs (no page content), let's break down how to approach this with both tools, plus answer your key question clearly.

Core Question: Do I need to download full pages?

Short answer: Absolutely not. You don’t have to fetch entire page bodies to validate URLs or collect internal links. Using HEAD requests instead of GET lets you only retrieve HTTP headers (including status codes like 200) without downloading the full content—this is way more efficient for your use case, saving bandwidth and time.


Using node-crawler for Internal URL Collection

node-crawler gives you granular control over request types and URL filtering. Here’s how to set it up for your needs:

  1. Filter Internal URLs: Use the filter option to only process URLs that match your target domain.
  2. Use HEAD Requests: Configure the crawler to send HEAD requests instead of GET to skip downloading page content.
  3. Track Valid URLs: Check for 200 status codes, then add valid internal URLs to your collection.

Quick Example Code

const Crawler = require('crawler');

const targetDomain = 'your-target-domain.com';
const collectedUrls = new Set();

const crawler = new Crawler({
  method: 'HEAD', // Send HEAD requests to avoid downloading content
  filter: (url) => url.hostname === targetDomain, // Only keep internal URLs
  callback: (error, res, done) => {
    if (error) {
      console.log(`Error with ${res.options.uri}: ${error}`);
    } else if (res.statusCode === 200) {
      collectedUrls.add(res.options.uri);
      console.log(`Added valid URL: ${res.options.uri}`);
      
      // Optional: If you need to extract links from pages (requires switching to GET)
      // Uncomment below and change method to GET if you need to scrape links from page bodies
      // $ = res.$;
      // $('a').each((i, el) => {
      //   const link = $(el).attr('href');
      //   crawler.queue(link);
      // });
    }
    done();
  }
});

// Start crawling from your seed URL
crawler.queue(`https://${targetDomain}`);

Using simplecrawler for Internal URL Collection

simplecrawler is built specifically for URL discovery, so it has built-in tools to limit crawling to specific domains. Here’s how to optimize it for your task:

  1. Restrict to Target Domain: Use domainWhitelist to ensure the crawler only processes URLs from your target domain.
  2. Enable HEAD Requests: Set requestMethod: 'HEAD' to skip downloading full page content.
  3. Capture Valid URLs: Listen for the fetchcomplete event to log URLs that return a 200 status code.

Quick Example Code

const Crawler = require('simplecrawler');

const targetDomain = 'your-target-domain.com';
const collectedUrls = new Set();

const crawler = new Crawler(`https://${targetDomain}`);
crawler.domainWhitelist = [targetDomain]; // Only crawl internal URLs
crawler.requestMethod = 'HEAD'; // Use HEAD requests to avoid content download

crawler.on('fetchcomplete', (queueItem, responseBuffer, response) => {
  if (response.statusCode === 200) {
    collectedUrls.add(queueItem.url);
    console.log(`Added valid URL: ${queueItem.url}`);
  }
});

crawler.on('complete', () => {
  console.log('Crawling finished! Collected URLs:');
  console.log([...collectedUrls]);
});

crawler.start();

Key Takeaways

  • Both tools let you skip full page downloads by using HEAD requests.
  • For internal URL filtering, simplecrawler’s domainWhitelist is more straightforward, while node-crawler gives you flexible control via the filter function.
  • Since you don’t need page content, sticking to HEAD requests will make your crawler faster and use far less bandwidth.

内容的提问来源于stack exchange,提问作者Machiavelli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:58:09