node-crawler与simplecrawler技术咨询:仅爬取内部URL相关疑问
Hey there! Since you're new to Node.js web crawlers and only need to collect internal URLs (no page content), let's break down how to approach this with both tools, plus answer your key question clearly.
Core Question: Do I need to download full pages?
Short answer: Absolutely not. You don’t have to fetch entire page bodies to validate URLs or collect internal links. Using HEAD requests instead of GET lets you only retrieve HTTP headers (including status codes like 200) without downloading the full content—this is way more efficient for your use case, saving bandwidth and time.
Using node-crawler for Internal URL Collection
node-crawler gives you granular control over request types and URL filtering. Here’s how to set it up for your needs:
- Filter Internal URLs: Use the
filteroption to only process URLs that match your target domain. - Use HEAD Requests: Configure the crawler to send
HEADrequests instead ofGETto skip downloading page content. - Track Valid URLs: Check for 200 status codes, then add valid internal URLs to your collection.
Quick Example Code
const Crawler = require('crawler'); const targetDomain = 'your-target-domain.com'; const collectedUrls = new Set(); const crawler = new Crawler({ method: 'HEAD', // Send HEAD requests to avoid downloading content filter: (url) => url.hostname === targetDomain, // Only keep internal URLs callback: (error, res, done) => { if (error) { console.log(`Error with ${res.options.uri}: ${error}`); } else if (res.statusCode === 200) { collectedUrls.add(res.options.uri); console.log(`Added valid URL: ${res.options.uri}`); // Optional: If you need to extract links from pages (requires switching to GET) // Uncomment below and change method to GET if you need to scrape links from page bodies // $ = res.$; // $('a').each((i, el) => { // const link = $(el).attr('href'); // crawler.queue(link); // }); } done(); } }); // Start crawling from your seed URL crawler.queue(`https://${targetDomain}`);
Using simplecrawler for Internal URL Collection
simplecrawler is built specifically for URL discovery, so it has built-in tools to limit crawling to specific domains. Here’s how to optimize it for your task:
- Restrict to Target Domain: Use
domainWhitelistto ensure the crawler only processes URLs from your target domain. - Enable HEAD Requests: Set
requestMethod: 'HEAD'to skip downloading full page content. - Capture Valid URLs: Listen for the
fetchcompleteevent to log URLs that return a 200 status code.
Quick Example Code
const Crawler = require('simplecrawler'); const targetDomain = 'your-target-domain.com'; const collectedUrls = new Set(); const crawler = new Crawler(`https://${targetDomain}`); crawler.domainWhitelist = [targetDomain]; // Only crawl internal URLs crawler.requestMethod = 'HEAD'; // Use HEAD requests to avoid content download crawler.on('fetchcomplete', (queueItem, responseBuffer, response) => { if (response.statusCode === 200) { collectedUrls.add(queueItem.url); console.log(`Added valid URL: ${queueItem.url}`); } }); crawler.on('complete', () => { console.log('Crawling finished! Collected URLs:'); console.log([...collectedUrls]); }); crawler.start();
Key Takeaways
- Both tools let you skip full page downloads by using
HEADrequests. - For internal URL filtering, simplecrawler’s
domainWhitelistis more straightforward, while node-crawler gives you flexible control via thefilterfunction. - Since you don’t need page content, sticking to
HEADrequests will make your crawler faster and use far less bandwidth.
内容的提问来源于stack exchange,提问作者Machiavelli

