You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Node.js从GCS存储桶下载、解压并读取CSV文件

问题描述

我正尝试使用Node.js(Google Cloud Functions)从Google Cloud Storage(GCS)存储桶读取CSV文件,该CSV文件为压缩包形式。现有实现代码如下:

const {google} = require('googleapis');
const {Storage} = require('@google-cloud/storage');
const {parse} = require("csv-parse");

const storage = new Storage({
  projectId: 'XXX',
  keyFilename: './api-project-XXX.json'
});

const file = storage.bucket('pubsite_prod_rev_XXX').file("sales/salesreport_202406.zip");

const content = await file.download({destination: "/tmp/test"}, function(err, file) {
  if (err) {
    console.log(err);
  } else {
    console.log("download successful");
  }
});

请问如何实现解压并读取该CSV文件?是否应使用createReadStream?是否需要先将文件存回存储桶,还是可以直接解压?


解决方案

核心结论

  • 不需要将文件存回GCS存储桶,可直接在Cloud Functions的/tmp临时目录完成解压与读取
  • 推荐使用createReadStream配合流式解压库处理,避免大文件一次性加载导致内存溢出

具体实现

1. 安装依赖

添加unzipper库(支持流式解压,适合处理大文件):

npm install unzipper

2. 流式处理(推荐)

直接通过流读取ZIP文件、解压并解析CSV,无需落地整个文件:

const {Storage} = require('@google-cloud/storage');
const {parse} = require("csv-parse");
const unzipper = require('unzipper');

const storage = new Storage({
  projectId: 'XXX',
  keyFilename: './api-project-XXX.json'
});

async function processZippedCSV() {
  const zipFile = storage.bucket('pubsite_prod_rev_XXX').file("sales/salesreport_202406.zip");

  // 流式读取→解压→解析CSV
  const stream = zipFile.createReadStream()
    .pipe(unzipper.Parse())
    .on('entry', (entry) => {
      // 仅处理ZIP内的CSV文件
      if (entry.path.endsWith('.csv')) {
        entry.pipe(parse({
          delimiter: ',',
          columns: true // 将第一行作为表头
        }))
        .on('data', (row) => {
          // 处理单条CSV数据,如打印、入库等
          console.log('CSV行数据:', row);
        })
        .on('end', () => {
          console.log('当前CSV文件解析完成');
          entry.autodrain(); // 释放资源
        })
        .on('error', (err) => {
          console.error('CSV解析错误:', err);
          entry.autodrain();
        });
      } else {
        entry.autodrain(); // 跳过非CSV文件
      }
    })
    .on('error', (err) => {
      console.error('解压错误:', err);
    });

  // 等待流式处理完成
  await new Promise((resolve, reject) => {
    stream.on('finish', resolve);
    stream.on('error', reject);
  });
}

// Cloud Functions入口根据触发器调整,HTTP触发器需导出函数
processZippedCSV().catch(console.error);

3. 小文件替代方案:下载后解压

如果ZIP文件体积很小,也可以先下载到/tmp再解压读取:

const fs = require('fs').promises;
const unzipper = require('unzipper');

async function downloadAndProcess() {
  const zipPath = '/tmp/salesreport_202406.zip';
  // 下载ZIP到临时目录
  await storage.bucket('pubsite_prod_rev_XXX').file("sales/salesreport_202406.zip").download({
    destination: zipPath
  });

  // 解压到指定目录
  const unzipDir = '/tmp/unzipped';
  await fs.mkdir(unzipDir, { recursive: true });
  await unzipper.Open.file(zipPath).then(d => d.extract({ path: unzipDir }));

  // 找到并读取CSV文件
  const files = await fs.readdir(unzipDir);
  const csvFile = files.find(f => f.endsWith('.csv'));
  if (csvFile) {
    const csvContent = await fs.readFile(`${unzipDir}/${csvFile}`, 'utf8');
    parse(csvContent, { columns: true }, (err, rows) => {
      if (err) console.error(err);
      else console.log('CSV数据:', rows);
    });
  }
}

关键说明

  • createReadStream的优势:对于大体积的ZIP/CSV文件,流式处理不会将整个文件加载到内存,避免触发Cloud Functions的内存配额限制
  • 临时目录/tmp:Cloud Functions提供最大512MB的临时存储空间,足够处理绝大多数压缩文件,处理完成后无需手动清理(函数结束后会自动回收)
  • 权限要求:确保Cloud Functions的服务账号拥有GCS存储桶的storage.objects.get权限,能正常读取目标ZIP文件

内容的提问来源于stack exchange,提问作者desmeit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 18:40:19