You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多线程下载文件/视频时文件体积异常增大的问题求助

多线程下载文件体积过大问题排查与修复

问题原因

  1. 服务器分片支持未验证:未检查服务器是否支持Range分片请求(通过响应头Accept-Ranges判断),若服务器不支持,每个线程都会下载完整文件,多次写入导致文件体积变为原文件的N倍(N为线程数)
  2. 文件写入逻辑错误:使用ab追加模式写入,当分片请求失效时,所有线程的完整文件内容都会被追加到同一文件中,直接导致体积膨胀
  3. 进度条线程安全问题:多线程共享tqdm进度条,可能出现重复更新,导致进度显示异常(虽不直接影响文件体积,但属于潜在问题)

修复后的代码

import os
import requests
from tqdm import tqdm
import threading
from requests.exceptions import RequestException

def download_chunk(url, save_path, start_byte, end_byte, progress_bar, lock):
    headers = {'Range': f'bytes={start_byte}-{end_byte}'}
    try:
        response = requests.get(url, headers=headers, stream=True, allow_redirects=True)
        response.raise_for_status()  # 检查请求是否成功

        with open(save_path, 'rb+') as file:
            file.seek(start_byte)
            for data in response.iter_content(chunk_size=4096):
                file.write(data)
                # 使用锁保证进度条更新线程安全
                with lock:
                    progress_bar.update(len(data))
    except RequestException as e:
        print(f"Chunk download failed: {e}")
        raise

def download_with_progress(url, save_folder, filename, num_threads=4):
    os.makedirs(save_folder, exist_ok=True)
    save_path = os.path.join(save_folder, filename)

    # 先验证服务器是否支持分片下载
    try:
        head_response = requests.head(url, allow_redirects=True)
        head_response.raise_for_status()
    except RequestException as e:
        print(f"Head request failed: {e}")
        return

    file_size = int(head_response.headers.get('content-length', 0))
    accept_ranges = head_response.headers.get('Accept-Ranges')

    if accept_ranges != 'bytes' or file_size == 0:
        # 不支持分片或无法获取文件大小,退化为单线程下载
        print("Server doesn't support range requests or file size unknown, using single thread download")
        progress_bar = tqdm(total=file_size, unit='B', unit_scale=True)
        try:
            response = requests.get(url, stream=True, allow_redirects=True)
            response.raise_for_status()
            with open(save_path, 'wb') as file:
                for data in response.iter_content(chunk_size=4096):
                    file.write(data)
                    progress_bar.update(len(data))
            progress_bar.close()
            print(f"Downloaded {filename}")
        except RequestException as e:
            print(f"Download failed: {e}")
        return

    # 支持分片下载,启动多线程
    block_size = file_size // num_threads
    progress_bar = tqdm(total=file_size, unit='B', unit_scale=True)
    lock = threading.Lock()  # 用于进度条线程安全更新

    # 预创建空文件并设置大小
    with open(save_path, 'wb') as file:
        file.truncate(file_size)

    threads = []
    for i in range(num_threads):
        start_byte = i * block_size
        end_byte = start_byte + block_size - 1 if i < num_threads - 1 else file_size - 1
        print(f"Thread {i}: Downloading bytes {start_byte}-{end_byte}")
        thread = threading.Thread(
            target=download_chunk,
            args=(url, save_path, start_byte, end_byte, progress_bar, lock)
        )
        threads.append(thread)
        thread.start()

    # 等待所有线程完成
    for thread in threads:
        thread.join()

    progress_bar.close()
    # 验证最终文件大小是否正确
    final_size = os.path.getsize(save_path)
    if final_size == file_size:
        print(f"Downloaded {filename} successfully, size matches original")
    else:
        print(f"Warning: Downloaded file size {final_size} doesn't match original {file_size}")

# Example usage:
download_with_progress("https://example.com/download.mp4", "GoogleColab", "hellyrev.mp4", num_threads=20)

关键修复点

  • 增加分片支持验证:通过head请求检查Accept-Ranges头,不支持则自动切换单线程
  • 修改文件写入方式:使用rb+模式并通过seek()定位到分片起始位置写入,避免追加导致的重复内容
  • 线程安全的进度条:添加线程锁Lock(),保证多线程更新进度条时不会出现混乱
  • 增加错误处理:捕获请求异常,避免程序崩溃
  • 文件大小验证:下载完成后对比最终文件大小与原文件大小,确保下载完整

内容的提问来源于stack exchange,提问作者Justahelper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 02:58:16