Node.js编写Webscraper遇DOM动态渲染问题求解决方案
Hey there, I’ve run into this exact problem countless times when scraping modern JS-heavy sites—tools like request only grab the initial, server-sent HTML, which for React/Angular/jQuery sites is usually just a root container and references to JS bundles, not the fully rendered content you’re after. Here are the most reliable solutions I’ve used to get around this:
1. Use a Headless Browser (Best for Full DOM Rendering)
Headless browsers simulate a real user’s browser session: they load all JS bundles, execute client-side code, and render the complete DOM just like Chrome/Firefox would. This is the go-to approach when you need to scrape content that’s only generated after JS runs.
My top picks are Puppeteer (Chrome-based) and Playwright (supports Chrome, Firefox, Safari). Here’s a quick Puppeteer example to scrape leads:
First, install the package:
npm install puppeteer
Then write your scraper:
const puppeteer = require('puppeteer'); async function scrapeLeadData() { // Launch a headless browser (add `headless: false` to see the browser window) const browser = await puppeteer.launch({ headless: 'new' }); const page = await browser.newPage(); // Navigate to the target page and wait for JS to finish rendering await page.goto('https://your-target-site.com/leads', { waitUntil: 'networkidle2' // Waits for most network activity to stop }); // Wait for the specific lead elements to appear (avoids race conditions) await page.waitForSelector('.lead-card'); // Extract data directly from the rendered DOM const leads = await page.evaluate(() => { return Array.from(document.querySelectorAll('.lead-card')).map(card => ({ name: card.querySelector('.lead-name').textContent.trim(), company: card.querySelector('.lead-company').textContent.trim(), contact: card.querySelector('.lead-email').textContent.trim() })); }); await browser.close(); console.log('Scraped leads:', leads); } scrapeLeadData();
Pro tip: If the site detects headless browsers, try adding a custom user agent or disabling sandbox mode (add args: ['--no-sandbox'] to puppeteer.launch()).
2. Bypass the DOM—Scrape the Site’s API Directly
Most modern JS sites fetch data from backend APIs and then render it to the DOM. Instead of scraping the rendered HTML, you can often call these APIs directly for cleaner, faster data.
Here’s how to do it:
- Open your browser’s DevTools (F12) and go to the Network tab.
- Refresh the page and filter for
XHRorFetchrequests. - Look for requests that return JSON data matching the leads you want (check the response preview).
- Copy that API endpoint, then use a modern HTTP client like
node-fetchoraxiosto call it.
Example with node-fetch:
npm install node-fetch
const fetch = require('node-fetch'); async function fetchLeadsViaAPI() { const apiUrl = 'https://your-target-site.com/api/v1/public/leads'; const response = await fetch(apiUrl, { headers: { // Mimic a real browser request to avoid being blocked 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'application/json' } }); if (!response.ok) throw new Error(`API request failed: ${response.status}`); const leads = await response.json(); console.log('Fetched leads from API:', leads); } fetchLeadsViaAPI();
Note: Some APIs require authentication (like cookies or bearer tokens). You can copy these from the DevTools request headers and include them in your fetch call.
3. Check for Server-Side Rendered (SSR) Content
A small number of React/Angular sites use SSR (e.g., Next.js for React, Nuxt.js for Vue), which means the initial HTML already includes rendered content. In this case, you might be able to use cheerio directly—just make sure to check the raw HTML first (use curl or node-fetch to get it) to see if your target data is present.
If it is, you can skip the headless browser and use a setup like this:
const fetch = require('node-fetch'); const cheerio = require('cheerio'); async function scrapeSSRLeads() { const response = await fetch('https://ssr-target-site.com/leads'); const html = await response.text(); const $ = cheerio.load(html); const leads = $('.lead-item').map((i, el) => ({ name: $(el).find('.lead-name').text().trim(), phone: $(el).find('.lead-phone').text().trim() })).get(); console.log('Scraped SSR leads:', leads); } scrapeSSRLeads();
Quick Note: Replace request (It’s Deprecated!)
The request package has been deprecated for years. For static HTTP requests, switch to node-fetch, axios, or got—they’re actively maintained and have better features.
内容的提问来源于stack exchange,提问作者z0idb3rg

