You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何跟踪Ruby中IO.copy_stream下载URL文件的进度?

跟踪大文件下载进度的实现方案

刚好做过类似的需求,给你几个实用的实现思路,既能跟踪下载进度,又能尽量保持下载效率:

方法一:自定义进度跟踪流(推荐,兼容IO.copy_stream)

这个方法通过包装输入流,在每次读取数据时统计已下载字节数,配合提前获取的文件总大小计算进度百分比。

步骤1:获取远程文件总大小

首先通过HEAD请求获取文件的Content-Length(注意:部分服务器可能不返回这个字段,需要做兼容处理):

require 'net/http'
require 'uri'

url = URI.parse("https://aws-file-url.com/bucket/large-file.mov")
total_size = 0

Net::HTTP.start(url.host, url.port, use_ssl: url.scheme == 'https') do |http|
  response = http.head(url.path)
  total_size = response['Content-Length'].to_i if response['Content-Length']
end

步骤2:实现进度跟踪IO类

这个类会包装原始输入流,每次读取数据时触发进度回调:

class ProgressTrackingIO < IO
  def initialize(io, total_size, &progress_callback)
    @io = io
    @total_size = total_size
    @downloaded = 0
    @progress_callback = progress_callback
  end

  def read(length = nil, outbuf = nil)
    data = @io.read(length, outbuf)
    if data
      @downloaded += data.size
      # 计算进度百分比(如果能获取总大小的话)
      progress = @total_size > 0 ? (@downloaded.to_f / @total_size) * 100 : nil
      @progress_callback.call(progress, @downloaded, @total_size) if @progress_callback
    end
    data
  end
end

步骤3:结合IO.copy_stream使用

用包装后的流替代原始URL,同时传入进度回调逻辑:

path = Rails.root.join("tmp", "#{SecureRandom.hex(12)}#{Time.now.to_i}")

Net::HTTP.start(url.host, url.port, use_ssl: url.scheme == 'https') do |http|
  http.get(url.path) do |response|
    tracking_io = ProgressTrackingIO.new(response, total_size) do |progress, downloaded, total|
      if progress
        puts "下载进度: #{progress.round(2)}% (#{downloaded}/#{total} bytes)"
      else
        puts "已下载: #{downloaded} bytes(无法获取总大小)"
      end
      # 这里可以替换成你的业务逻辑:比如写入日志、更新数据库状态、推送前端通知等
    end
    IO.copy_stream(tracking_io, path)
  end
end

方法二:分块读写手动统计进度

如果不想自定义IO类,也可以直接分块读取远程数据,写入本地文件时实时统计进度:

require 'net/http'
require 'uri'

url = URI.parse("https://aws-file-url.com/bucket/large-file.mov")
path = Rails.root.join("tmp", "#{SecureRandom.hex(12)}#{Time.now.to_i}")
chunk_size = 1024 * 1024 # 每次读取1MB块,可根据需求调整

Net::HTTP.start(url.host, url.port, use_ssl: url.scheme == 'https') do |http|
  http.get(url.path) do |response|
    total_size = response['Content-Length'].to_i if response['Content-Length']
    downloaded = 0

    File.open(path, 'wb') do |file|
      response.read_body do |chunk|
        file.write(chunk)
        downloaded += chunk.size
        
        if total_size > 0
          progress = (downloaded.to_f / total_size) * 100
          puts "下载进度: #{progress.round(2)}% (#{downloaded}/#{total_size} bytes)"
        else
          puts "已下载: #{downloaded} bytes"
        end
      end
    end
  end
end

额外提示:如果是AWS S3文件

如果你的文件是存储在S3上的,直接使用AWS SDK for Ruby会更方便,SDK原生支持进度回调:

require 'aws-sdk-s3'

s3 = Aws::S3::Client.new(region: 'your-region')
bucket_name = 'your-bucket'
file_key = 'large-file.mov'
path = Rails.root.join("tmp", "#{SecureRandom.hex(12)}#{Time.now.to_i}")

# 获取文件总大小
total_size = s3.head_object(bucket: bucket_name, key: file_key).content_length
downloaded = 0

s3.get_object(bucket: bucket_name, key: file_key) do |chunk|
  File.open(path, 'ab') { |file| file.write(chunk) }
  downloaded += chunk.size
  progress = (downloaded.to_f / total_size) * 100
  puts "下载进度: #{progress.round(2)}%"
end

注意事项

  • 部分服务器可能不会返回Content-Length(比如动态生成的文件),这时候只能显示已下载字节数,无法计算百分比。
  • 网络中断等异常情况需要额外捕获处理,避免下载失败后留下不完整文件。
  • 自定义IO类的方法几乎不会影响下载效率,因为只是在读取数据时做了简单的统计逻辑。

内容的提问来源于stack exchange,提问作者Graham Slick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 02:30:23