You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Node.js的crawler爬取网站全部内部链接并获取页面标题

Hey there! Let's build out a working crawler using the crawler npm package that scrapes all internal links of a target website and captures each page's title. I'll expand on the sample code you found and fill in the critical missing pieces to make this work reliably.

Step 1: Install Dependencies

First, make sure you have the required packages installed. We'll use crawler for the crawling logic and Node's built-in url module to handle link resolution:

npm install crawler

Step 2: Full Working Code

Here's a complete implementation that tracks internal links, avoids duplicate crawls, and extracts page titles:

const Crawler = require("crawler");
const url = require("url");

// Define your target website's base URL
const BASE_URL = "https://your-target-website.com";
// Keep track of URLs we've already crawled to avoid loops and duplicates
const visitedUrls = new Set();

const c = new Crawler({
  maxConnections: 10,
  // Optional: Add a delay between requests to be respectful to the server
  rateLimit: 1000,
  callback: function (error, res, done) {
    if (error) {
      console.error(`Error crawling ${res.options.uri}:`, error);
      return done();
    }

    const $ = res.$;
    const currentUrl = res.options.uri;
    
    // Mark this URL as visited
    visitedUrls.add(currentUrl);
    
    // Extract and log the page title
    const pageTitle = $("title").text().trim() || "No title found";
    console.log(`Title: ${pageTitle} | URL: ${currentUrl}`);

    // Extract all links from the page
    $("a").each((index, element) => {
      let link = $(element).attr("href");
      if (!link) return;

      // Resolve relative links to absolute URLs
      const absoluteUrl = url.resolve(BASE_URL, link);
      
      // Check if this is an internal link (same domain as BASE_URL) and not visited yet
      const isInternal = url.parse(absoluteUrl).hostname === url.parse(BASE_URL).hostname;
      if (isInternal && !visitedUrls.has(absoluteUrl)) {
        // Add the internal link to the crawler queue
        c.queue(absoluteUrl);
        // Mark it as visited immediately to avoid duplicate queue entries
        visitedUrls.add(absoluteUrl);
      }
    });

    done();
  }
});

// Start crawling from the base URL
c.queue(BASE_URL);

Key Details Explained

  • visitedUrls Set: This prevents us from crawling the same page multiple times, which avoids infinite loops (like pages linking back to each other) and reduces unnecessary requests.
  • Link Resolution: Using url.resolve() converts relative paths (like /about or contact.html) into full absolute URLs, making it easier to validate and track them.
  • Internal Link Check: We compare the hostname of the extracted link with our base URL's hostname to ensure we only crawl pages within the target website.
  • Rate Limiting: The rateLimit option adds a 1-second delay between requests—this is important to be respectful of the target server and avoid getting blocked.

Important Notes

  • Always check the target website's robots.txt file (e.g., https://your-target-website.com/robots.txt) to make sure crawling is allowed.
  • Some websites may have anti-scraping measures (like CAPTCHAs or IP blocking). If you run into issues, consider adding user-agent headers or using proxies.
  • For larger websites, you might want to add more robust error handling (like retries for failed requests) or persist crawled data to a database instead of just logging it.

内容的提问来源于stack exchange,提问作者Alexander Solonik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:41:40