基于Request-Promise与Cheerio的Promise循环嵌套爬虫实现问题
Got it, let's walk through how to build this web scraper using Request-Promise and Cheerio step by step. Here's a complete implementation that follows your exact workflow:
Step-by-Step Scraper Implementation
First, make sure you have the required dependencies installed:
npm install request-promise cheerio
Then, here's the full code with detailed explanations:
const rp = require('request-promise'); const cheerio = require('cheerio'); // Initialize the empty array to store our structured data const scrapedData = []; // Create a cookie jar to persist login sessions (critical for authenticated requests) const cookieJar = rp.jar(); async function runScraper() { try { // --- Step 1: Complete Login --- const loginOptions = { uri: 'https://your-login-url.com/login', // Replace with your actual login endpoint method: 'POST', form: { username: 'your-username', // Add your login credentials here password: 'your-password' }, jar: cookieJar, // Save login cookies to reuse in subsequent requests resolveWithFullResponse: true, // Optional: Check login success via response simple: false // Don't throw errors for non-200 status (handles login redirects) }; const loginResponse = await rp(loginOptions); // Verify login worked (adjust check based on your site's response) if (loginResponse.statusCode === 200 || loginResponse.body.includes('Welcome')) { console.log('Login successful! Proceeding to scrape.'); } else { throw new Error('Login failed. Double-check credentials or login endpoint.'); } // --- Step 2: Scrape data from two target pages --- // Scrape Page 1 const page1Options = { uri: 'https://your-target-page-1.com', // Replace with first page URL jar: cookieJar, // Use authenticated cookies transform: (body) => cheerio.load(body) // Auto-parse HTML with Cheerio }; const $page1 = await rp(page1Options); // Extract links from Page 1 (update selector to match your site's elements) $page1('.target-link-class').each((index, element) => { const link = $page1(element).attr('href'); scrapedData.push({ link: link, items: [] }); }); // Scrape Page 2 (repeat similar logic for the second page) const page2Options = { uri: 'https://your-target-page-2.com', // Replace with second page URL jar: cookieJar, transform: (body) => cheerio.load(body) }; const $page2 = await rp(page2Options); $page2('.target-link-class').each((index, element) => { const link = $page2(element).attr('href'); scrapedData.push({ link: link, items: [] }); }); console.log(`Collected ${scrapedData.length} links to process next.`); // --- Step 3: Process each link in the scrapedData array --- for (const entry of scrapedData) { console.log(`Processing link: ${entry.link}`); const detailPageOptions = { uri: entry.link, jar: cookieJar, transform: (body) => cheerio.load(body) }; const $detailPage = await rp(detailPageOptions); // Extract items from the detail page (update selector to match your target items) $detailPage('.item-element-class').each((index, element) => { const itemContent = $detailPage(element).text().trim(); entry.items.push(itemContent); }); console.log(`Added ${entry.items.length} items for this link.`); } // Final output: Use or log the fully populated scrapedData array console.log('Scraping complete! Final structured data:', scrapedData); } catch (error) { console.error('Scraper hit an issue:', error.message); } } // Kick off the scraper runScraper();
Key Details to Adjust:
- Selectors: Replace
.target-link-classand.item-element-classwith actual CSS selectors from your target sites. Use your browser's dev tools to inspect elements and get these selectors. - URLs: Swap out all placeholder URLs (login, target pages) with the real ones you're working with.
- Login Verification: Adjust the login success check to match how your site confirms a successful login (e.g., checking for a specific cookie or page content).
Pro Tips:
- Add small delays between requests (wrap
setTimeoutin a promise) to avoid triggering rate limits or getting blocked by the site. - Resolve relative URLs to absolute ones if your target links aren't full URLs (use
new URL(link, baseUrl)). - Add more robust error handling for individual requests (e.g., skip broken links instead of crashing the whole scraper).
内容的提问来源于stack exchange,提问作者Chris Talke
相关产品推荐
相关产品推荐

