You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Puppeteer批量抓取books.toscrape.com图书标题、价格及库存信息的代码故障排查求助

Fixing Puppeteer + Cheerio Bulk Scraping for Books to Scrape

Let's break down what's going wrong with your code first, then walk through a fully working solution.

Issues in Your Current Code

Here are the key problems preventing your script from running properly:

  • Undefined url variable: You reference url inside the async function but never assign it to objArray[i].Url. Also, your URLs are relative—you need to prepend http:// to make them valid full URLs.
  • Broken CSS selectors:
    • The stock class is instock availability (with a space), not instockavailability as you wrote. Use p.instock.availability to target it correctly.
    • The price selector p[price_color] is invalid; class selectors need a dot, so it should be p.price_color.
  • Inefficient browser usage: Launching a new browser for every URL is slow and wastes resources. You should launch one browser instance and reuse it for all pages.
  • Overwriting output: fs.writeFileSync replaces the file every loop iteration, so you'll only ever get data from the last book. Collect all results first, then write once.
  • Unclean data extraction:
    • Grabbing the entire product_main div's text for the title includes extra content (price, stock labels, etc.). Target the <h1> tag inside that div instead.
    • The stock text has messy whitespace—use trim() to clean it up.
  • Missing error handling & cleanup: No catch block to handle crashes, and you never close the browser, leaving orphaned processes running in the background.

Corrected Implementation

Here's the fixed script that addresses all these issues:

const puppeteer = require('puppeteer');
const cheerio = require('cheerio');
const fs = require('fs');

// Updated URL array with full valid URLs
const objArray = [ 
  {"Url": "http://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"}, 
  {"Url": "http://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html"}, 
  {"Url": "http://books.toscrape.com/catalogue/soumission_998/index.html"} 
];

(async () => {
  let browser;
  try {
    // Launch browser ONCE before processing all URLs
    browser = await puppeteer.launch({ headless: false });
    const allProductDetails = [];

    for (const book of objArray) {
      const page = await browser.newPage();
      // Wait for DOM to fully load before scraping
      await page.goto(book.Url, { waitUntil: 'domcontentloaded' });

      const content = await page.content();
      const $ = cheerio.load(content);
      const productMain = $('.col-sm-6.product_main');

      // Clean and extract precise data
      const title = productMain.find('h1').text().trim();
      const price = productMain.find('p.price_color').text().trim();
      const stock = productMain.find('p.instock.availability').text().trim();

      allProductDetails.push({
        Title: title,
        Price: price,
        Stock: stock
      });

      await page.close(); // Close individual page to free up memory
    }

    // Write all collected data to file in one go
    fs.writeFileSync('files.json', JSON.stringify(allProductDetails, null, 2), 'utf8');
    console.log('Scraping complete! Data saved to files.json');
    console.log(allProductDetails);

  } catch (error) {
    console.error('Something went wrong during scraping:', error);
  } finally {
    // Always close the browser, even if an error occurs
    if (browser) await browser.close();
  }
})();

Key Improvements Explained

  • Single browser instance: Reusing one browser drastically speeds up scraping and reduces resource usage.
  • Valid URLs: Fixed relative paths to full working URLs.
  • Precise selectors: Targeted exact elements (h1 for title, correct class selectors for price/stock) to avoid irrelevant data.
  • Clean data: Used trim() to remove extra newlines and whitespace from extracted text.
  • Batch writing: Collected all results first, then wrote to the file once to prevent overwriting.
  • Robust error handling: Added a catch block to log issues, and a finally block to ensure the browser always closes.
  • DOM load wait: Added waitUntil: 'domcontentloaded' to make sure the page content is fully loaded before attempting to scrape.

内容的提问来源于stack exchange,提问作者Divyanshu Sah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 18:12:40