You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Puppeteer编写爬虫时出现ReferenceError: fs is not defined错误求助

解决Puppeteer爬虫中fs is not defined错误

问题原因

你在page.evaluate()的回调里调用了fs模块,但这个回调是运行在浏览器前端环境中的,而fs是Node.js专属的文件系统模块,前端环境根本不存在这个变量,所以会抛出ReferenceError。

解决方案

把文件操作、图片下载这类Node.js环境的逻辑,从浏览器上下文移到主进程中;page.evaluate()只负责从页面提取所需的数据(比如品牌标题、logo链接),然后把数据返回给主进程处理。

修改后的完整代码:

const puppeteer = require('puppeteer');
const fs = require('fs/promises');
const path = require('path'); // 用path处理路径更可靠

async function scrapeWebsite() {
    const browser = await puppeteer.launch({ headless: false });
    const page = await browser.newPage();

    try {
        await page.goto('url of the website');
        await page.click('div[role="button"]');

        // 从浏览器上下文提取数据,返回给主进程
        const brandList = await page.evaluate(async () => {
            let count = 0;
            const brands = [];
            while (count < 1) {
                const items = document.querySelectorAll('ul li a[title]');
                for (const item of items) {
                    count++;
                    const brand_element = item.querySelector('img');
                    brands.push({
                        title: brand_element.getAttribute('title'),
                        logoUrl: brand_element.getAttribute('src')
                    });
                    if (count >= 1) break;
                }
            }
            return brands;
        });

        // 主进程中处理文件操作和图片下载
        const currentDirectory = 'C:/Users/xxx/Documents/Project/xxx/Assets';
        for (const brand of brandList) {
            const folderPath = path.join(currentDirectory, brand.title); // 用path.join避免路径拼接错误
            await fs.mkdir(folderPath, { recursive: true });
            
            // 下载图片:用page.goto获取图片二进制数据
            const imageResponse = await page.goto(brand.logoUrl);
            const imageBuffer = await imageResponse.buffer();
            
            const imagePath = path.join(folderPath, `${brand.title}.jpg`);
            await fs.writeFile(imagePath, imageBuffer);
        }

    } catch (error) {
        console.error('Error:', error);
    } finally {
        await browser.close(); // 确保爬虫结束后关闭浏览器
    }
}

scrapeWebsite();

关键修改点

  • 将fs相关操作全部移出page.evaluate(),放到主进程执行
  • 用page.evaluate()提取页面数据后返回,主进程再处理文件和下载
  • 使用path.join()处理路径,避免手动拼接出现的斜杠问题
  • 补充了图片二进制数据的下载逻辑(原代码仅写入空文件)
  • 添加finally块确保浏览器正常关闭

内容的提问来源于stack exchange,提问作者Ranjith Varatharajan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 16:35:20