使用Cheerio.js爬取网页时遇问题:仅能获取单个文章链接
Hey there! I’ve been in your shoes before—frustrated when web scraping only pulls a single link instead of all the ones you need. Let’s walk through the most common fixes for this issue with Node.js, Cheerio, and Request.
1. You’re Grabbing a Single Element Instead of Iterating
The #1 mistake here is using a method that returns just one element (like .attr()) instead of looping through all matching elements. For example, if you write:
const link = $('a.article-link').attr('href'); console.log(link); // Only logs the first link
This will only get the href of the first matching <a> tag. Instead, use Cheerio’s .each() method to loop through every matching element:
const request = require('request'); const cheerio = require('cheerio'); request('https://your-target-site.com', (error, response, html) => { if (!error && response.statusCode === 200) { const $ = cheerio.load(html); // Iterate over all article links $('a.article-link').each((index, element) => { const link = $(element).attr('href'); console.log(`Article Link ${index + 1}: ${link}`); }); } });
2. Your Selector Is Too Specific (or Not Specific Enough)
Double-check that your CSS selector is matching all the links you want. Sometimes a class name is only used on the first article, or you’re targeting a parent element that only appears once.
- Open your target site in Chrome/Firefox, right-click an article link, and select Inspect
- Right-click the element in DevTools and choose Copy > Copy selector
- Paste that selector into your Cheerio code to ensure it’s targeting all relevant links
For example, if the links are inside <div class="article-card">, your selector should be div.article-card a instead of just a.article-link (if that class isn’t applied to all links).
3. The Page Uses Dynamic Content (AJAX/JavaScript)
Request only fetches the static HTML of the page. If the article links are loaded dynamically with JavaScript (like after an API call), Cheerio won’t see them—since it doesn’t execute JS.
If this is the case, you’ll need to switch to a headless browser tool like Puppeteer instead of Request. Here’s a quick snippet to show how that works:
const puppeteer = require('puppeteer'); (async () => { const browser = await puppeteer.launch(); const page = await browser.newPage(); await page.goto('https://your-target-site.com'); const links = await page.$$eval('a.article-link', elements => elements.map(el => el.href)); console.log(links); // Logs all dynamically loaded links await browser.close(); })();
4. You’re Not Handling Asynchronous Code Correctly
If you’re trying to collect links into an array and use them outside the request callback, make sure you’re handling the async flow properly. Using request-promise-native (a Promise-based version of Request) can make this easier:
const request = require('request-promise-native'); const cheerio = require('cheerio'); async function scrapeLinks() { const html = await request('https://your-target-site.com'); const $ = cheerio.load(html); const links = []; $('a.article-link').each((index, element) => { links.push($(element).attr('href')); }); return links; } scrapeLinks().then(links => console.log(links)).catch(err => console.error(err));
Give these fixes a try—9 times out of 10, it’s either a missing loop or a bad selector. Let me know if you’re still stuck!
内容的提问来源于stack exchange,提问作者Jay Patel

