You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Request-Promise与Cheerio的Promise循环嵌套爬虫实现问题

Got it, let's walk through how to build this web scraper using Request-Promise and Cheerio step by step. Here's a complete implementation that follows your exact workflow:

Step-by-Step Scraper Implementation

First, make sure you have the required dependencies installed:

npm install request-promise cheerio

Then, here's the full code with detailed explanations:

const rp = require('request-promise');
const cheerio = require('cheerio');

// Initialize the empty array to store our structured data
const scrapedData = [];

// Create a cookie jar to persist login sessions (critical for authenticated requests)
const cookieJar = rp.jar();

async function runScraper() {
  try {
    // --- Step 1: Complete Login ---
    const loginOptions = {
      uri: 'https://your-login-url.com/login', // Replace with your actual login endpoint
      method: 'POST',
      form: {
        username: 'your-username', // Add your login credentials here
        password: 'your-password'
      },
      jar: cookieJar, // Save login cookies to reuse in subsequent requests
      resolveWithFullResponse: true, // Optional: Check login success via response
      simple: false // Don't throw errors for non-200 status (handles login redirects)
    };

    const loginResponse = await rp(loginOptions);
    // Verify login worked (adjust check based on your site's response)
    if (loginResponse.statusCode === 200 || loginResponse.body.includes('Welcome')) {
      console.log('Login successful! Proceeding to scrape.');
    } else {
      throw new Error('Login failed. Double-check credentials or login endpoint.');
    }

    // --- Step 2: Scrape data from two target pages ---
    // Scrape Page 1
    const page1Options = {
      uri: 'https://your-target-page-1.com', // Replace with first page URL
      jar: cookieJar, // Use authenticated cookies
      transform: (body) => cheerio.load(body) // Auto-parse HTML with Cheerio
    };

    const $page1 = await rp(page1Options);
    // Extract links from Page 1 (update selector to match your site's elements)
    $page1('.target-link-class').each((index, element) => {
      const link = $page1(element).attr('href');
      scrapedData.push({ link: link, items: [] });
    });

    // Scrape Page 2 (repeat similar logic for the second page)
    const page2Options = {
      uri: 'https://your-target-page-2.com', // Replace with second page URL
      jar: cookieJar,
      transform: (body) => cheerio.load(body)
    };

    const $page2 = await rp(page2Options);
    $page2('.target-link-class').each((index, element) => {
      const link = $page2(element).attr('href');
      scrapedData.push({ link: link, items: [] });
    });

    console.log(`Collected ${scrapedData.length} links to process next.`);

    // --- Step 3: Process each link in the scrapedData array ---
    for (const entry of scrapedData) {
      console.log(`Processing link: ${entry.link}`);
      const detailPageOptions = {
        uri: entry.link,
        jar: cookieJar,
        transform: (body) => cheerio.load(body)
      };

      const $detailPage = await rp(detailPageOptions);
      // Extract items from the detail page (update selector to match your target items)
      $detailPage('.item-element-class').each((index, element) => {
        const itemContent = $detailPage(element).text().trim();
        entry.items.push(itemContent);
      });

      console.log(`Added ${entry.items.length} items for this link.`);
    }

    // Final output: Use or log the fully populated scrapedData array
    console.log('Scraping complete! Final structured data:', scrapedData);

  } catch (error) {
    console.error('Scraper hit an issue:', error.message);
  }
}

// Kick off the scraper
runScraper();

Key Details to Adjust:

  • Selectors: Replace .target-link-class and .item-element-class with actual CSS selectors from your target sites. Use your browser's dev tools to inspect elements and get these selectors.
  • URLs: Swap out all placeholder URLs (login, target pages) with the real ones you're working with.
  • Login Verification: Adjust the login success check to match how your site confirms a successful login (e.g., checking for a specific cookie or page content).

Pro Tips:

  • Add small delays between requests (wrap setTimeout in a promise) to avoid triggering rate limits or getting blocked by the site.
  • Resolve relative URLs to absolute ones if your target links aren't full URLs (use new URL(link, baseUrl)).
  • Add more robust error handling for individual requests (e.g., skip broken links instead of crashing the whole scraper).

内容的提问来源于stack exchange,提问作者Chris Talke

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:33:45