使用Puppeteer批量抓取books.toscrape.com图书标题、价格及库存信息的代码故障排查求助
Fixing Puppeteer + Cheerio Bulk Scraping for Books to Scrape
Let's break down what's going wrong with your code first, then walk through a fully working solution.
Issues in Your Current Code
Here are the key problems preventing your script from running properly:
- Undefined
urlvariable: You referenceurlinside the async function but never assign it toobjArray[i].Url. Also, your URLs are relative—you need to prependhttp://to make them valid full URLs. - Broken CSS selectors:
- The stock class is
instock availability(with a space), notinstockavailabilityas you wrote. Usep.instock.availabilityto target it correctly. - The price selector
p[price_color]is invalid; class selectors need a dot, so it should bep.price_color.
- The stock class is
- Inefficient browser usage: Launching a new browser for every URL is slow and wastes resources. You should launch one browser instance and reuse it for all pages.
- Overwriting output:
fs.writeFileSyncreplaces the file every loop iteration, so you'll only ever get data from the last book. Collect all results first, then write once. - Unclean data extraction:
- Grabbing the entire
product_maindiv's text for the title includes extra content (price, stock labels, etc.). Target the<h1>tag inside that div instead. - The stock text has messy whitespace—use
trim()to clean it up.
- Grabbing the entire
- Missing error handling & cleanup: No
catchblock to handle crashes, and you never close the browser, leaving orphaned processes running in the background.
Corrected Implementation
Here's the fixed script that addresses all these issues:
const puppeteer = require('puppeteer'); const cheerio = require('cheerio'); const fs = require('fs'); // Updated URL array with full valid URLs const objArray = [ {"Url": "http://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"}, {"Url": "http://books.toscrape.com/catalogue/tipping-the-velvet_999/index.html"}, {"Url": "http://books.toscrape.com/catalogue/soumission_998/index.html"} ]; (async () => { let browser; try { // Launch browser ONCE before processing all URLs browser = await puppeteer.launch({ headless: false }); const allProductDetails = []; for (const book of objArray) { const page = await browser.newPage(); // Wait for DOM to fully load before scraping await page.goto(book.Url, { waitUntil: 'domcontentloaded' }); const content = await page.content(); const $ = cheerio.load(content); const productMain = $('.col-sm-6.product_main'); // Clean and extract precise data const title = productMain.find('h1').text().trim(); const price = productMain.find('p.price_color').text().trim(); const stock = productMain.find('p.instock.availability').text().trim(); allProductDetails.push({ Title: title, Price: price, Stock: stock }); await page.close(); // Close individual page to free up memory } // Write all collected data to file in one go fs.writeFileSync('files.json', JSON.stringify(allProductDetails, null, 2), 'utf8'); console.log('Scraping complete! Data saved to files.json'); console.log(allProductDetails); } catch (error) { console.error('Something went wrong during scraping:', error); } finally { // Always close the browser, even if an error occurs if (browser) await browser.close(); } })();
Key Improvements Explained
- Single browser instance: Reusing one browser drastically speeds up scraping and reduces resource usage.
- Valid URLs: Fixed relative paths to full working URLs.
- Precise selectors: Targeted exact elements (h1 for title, correct class selectors for price/stock) to avoid irrelevant data.
- Clean data: Used
trim()to remove extra newlines and whitespace from extracted text. - Batch writing: Collected all results first, then wrote to the file once to prevent overwriting.
- Robust error handling: Added a
catchblock to log issues, and afinallyblock to ensure the browser always closes. - DOM load wait: Added
waitUntil: 'domcontentloaded'to make sure the page content is fully loaded before attempting to scrape.
内容的提问来源于stack exchange,提问作者Divyanshu Sah
相关产品推荐
相关产品推荐

