You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Ruby 2.5和Rails 5.1流式下载并解压5GB ZIP文件?

实现边下载边解压大ZIP文件(Ruby 2.5 + Rails 5.1)

核心思路

咱们要做的是流式处理,完全不用把5GB的ZIP文件整个拉到内存里:一边用10KB的缓冲区流式下载远程文件,一边把实时拿到的数据传给解压工具,全程保持低内存占用。下面是具体的实现方案:

第一步:安装依赖

Ruby标准库没有ZIP处理能力,得先装rubyzip gem(选兼容Ruby 2.5的2.x版本)。在Rails项目的Gemfile里加一行:

gem 'rubyzip', '~> 2.3'

然后执行bundle install完成安装。

第二步:核心逻辑(用后台任务执行)

⚠️ 重要提醒:绝对不要在Rails控制器的请求周期里直接跑这段代码!5GB文件的处理时间很长,会直接导致请求超时。一定要用Rails的后台任务框架(比如Sidekiq、Delayed Job)来执行。

下面是完整的后台任务示例:

# app/jobs/unzip_downloaded_file_job.rb
class UnzipDownloadedFileJob < ApplicationJob
  queue_as :default

  def perform(zip_url, unzip_dir)
    require 'net/http'
    require 'zip'
    require 'io/pipe'
    require 'fileutils'

    BUFFER_SIZE = 10 * 1024 # 严格按照你要求的10KB缓冲区
    read_io, write_io = IO.pipe

    # 线程1:流式下载,把10KB块写入管道
    download_thread = Thread.new do
      begin
        uri = URI.parse(zip_url)
        # 自动适配HTTP/HTTPS请求
        Net::HTTP.start(uri.host, uri.port, use_ssl: uri.scheme == 'https') do |http|
          http.request_get(uri.path) do |response|
            # 先检查下载是否成功
            raise "下载失败,响应码:#{response.code}" unless response.is_a?(Net::HTTPSuccess)
            # 分块读取并写入管道
            response.read_body(BUFFER_SIZE) do |chunk|
              write_io.write(chunk)
            end
          end
        end
      rescue => e
        Rails.logger.error "下载出错:#{e.message}"
        raise e # 抛出异常让任务框架自动重试(可根据需求调整)
      ensure
        write_io.close # 写完关闭写端,通知解压线程结束
      end
    end

    # 线程2:从管道读数据,流式解压
    unzip_thread = Thread.new do
      begin
        Zip::InputStream.open(read_io) do |zis|
          while entry = zis.get_next_entry
            # 跳过ZIP内的目录,后续用mkdir_p自动创建
            next if entry.directory?

            # 构建目标文件路径,确保父目录存在
            target_path = File.join(unzip_dir, entry.name)
            FileUtils.mkdir_p(File.dirname(target_path))

            # 流式写入解压后的文件,同样用10KB缓冲区
            File.open(target_path, 'wb') do |f|
              while chunk = zis.read(BUFFER_SIZE)
                f.write(chunk)
              end
            end

            Rails.logger.info "已完成解压:#{entry.name}"
          end
        end
      rescue Zip::Error => e
        Rails.logger.error "ZIP文件损坏或解压失败:#{e.message}"
        # 可选:清理已解压的文件,避免残留
        FileUtils.rm_rf(unzip_dir) if File.exist?(unzip_dir)
        raise e
      rescue => e
        Rails.logger.error "解压流程出错:#{e.message}"
        raise e
      ensure
        read_io.close
      end
    end

    # 等待两个线程全部完成
    download_thread.join
    unzip_thread.join

    Rails.logger.info "所有文件已成功解压到:#{unzip_dir}"
  end
end

第三步:触发后台任务

在控制器里加一个接口来启动解压任务:

# app/controllers/unzip_controller.rb
class UnzipController < ApplicationController
  def start
    zip_url = params[:zip_url]
    # 生成唯一的解压目录,避免文件冲突
    unzip_dir = Rails.root.join('tmp', 'unzipped_files', SecureRandom.uuid).to_s

    # 启动后台任务
    UnzipDownloadedFileJob.perform_later(zip_url, unzip_dir)

    render json: {
      message: "解压任务已启动,文件将被解压到:#{unzip_dir}"
    }
  end
end

记得在routes.rb中添加路由:

post 'unzip/start', to: 'unzip#start'

关键注意事项

  1. 磁盘空间:提前确认服务器有足够的磁盘空间——5GB的ZIP文件解压后体积通常会更大;
  2. 后台任务配置:如果用Sidekiq,要确保Sidekiq worker进程处于运行状态;
  3. 异常处理:代码里已经加了基础的异常捕获和日志,你可以根据需求调整重试逻辑;
  4. HTTPS证书:如果下载地址是自签名HTTPS证书,可在Net::HTTP.start中添加verify_mode: OpenSSL::SSL::VERIFY_NONE(生产环境不推荐这么做);
  5. 进度监控:如果需要实时查看下载/解压进度,可以用Redis存储已下载字节数、已解压文件数,再做一个查询接口展示。

内容的提问来源于stack exchange,提问作者BraveVN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:57:15