使用Puppeteer编写爬虫时出现ReferenceError: fs is not defined错误求助
解决Puppeteer爬虫中
fs is not defined错误 问题原因
你在page.evaluate()的回调里调用了fs模块,但这个回调是运行在浏览器前端环境中的,而fs是Node.js专属的文件系统模块,前端环境根本不存在这个变量,所以会抛出ReferenceError。
解决方案
把文件操作、图片下载这类Node.js环境的逻辑,从浏览器上下文移到主进程中;page.evaluate()只负责从页面提取所需的数据(比如品牌标题、logo链接),然后把数据返回给主进程处理。
修改后的完整代码:
const puppeteer = require('puppeteer'); const fs = require('fs/promises'); const path = require('path'); // 用path处理路径更可靠 async function scrapeWebsite() { const browser = await puppeteer.launch({ headless: false }); const page = await browser.newPage(); try { await page.goto('url of the website'); await page.click('div[role="button"]'); // 从浏览器上下文提取数据,返回给主进程 const brandList = await page.evaluate(async () => { let count = 0; const brands = []; while (count < 1) { const items = document.querySelectorAll('ul li a[title]'); for (const item of items) { count++; const brand_element = item.querySelector('img'); brands.push({ title: brand_element.getAttribute('title'), logoUrl: brand_element.getAttribute('src') }); if (count >= 1) break; } } return brands; }); // 主进程中处理文件操作和图片下载 const currentDirectory = 'C:/Users/xxx/Documents/Project/xxx/Assets'; for (const brand of brandList) { const folderPath = path.join(currentDirectory, brand.title); // 用path.join避免路径拼接错误 await fs.mkdir(folderPath, { recursive: true }); // 下载图片:用page.goto获取图片二进制数据 const imageResponse = await page.goto(brand.logoUrl); const imageBuffer = await imageResponse.buffer(); const imagePath = path.join(folderPath, `${brand.title}.jpg`); await fs.writeFile(imagePath, imageBuffer); } } catch (error) { console.error('Error:', error); } finally { await browser.close(); // 确保爬虫结束后关闭浏览器 } } scrapeWebsite();
关键修改点
- 将
fs相关操作全部移出page.evaluate(),放到主进程执行 - 用
page.evaluate()提取页面数据后返回,主进程再处理文件和下载 - 使用
path.join()处理路径,避免手动拼接出现的斜杠问题 - 补充了图片二进制数据的下载逻辑(原代码仅写入空文件)
- 添加
finally块确保浏览器正常关闭
内容的提问来源于stack exchange,提问作者Ranjith Varatharajan
相关产品推荐
相关产品推荐

