You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多连接下载器暂停恢复后无限下载问题排查与修复方案咨询

Hey there, let’s work through this pause-resume issue with your Python multi-connection downloader. I’ve dealt with similar problems before, so let’s break this down step by step.

Fixing Your Multi-Connection Downloader’s Pause-Resume Failures

First, Let’s Diagnose the Core Issues

The progress exceeding 100%, negative remaining bytes, and IndexError you’re seeing are linked to two critical problems:

  1. Unclean connection termination on pause: When you pause, old connections aren’t properly shut down. When resuming, the server or local socket stack sends leftover bytes from those dead connections, which messes up your downloaded byte count and mmap writes.
  2. Thread-unsafe state tracking: Your progress and downloaded byte counters aren’t synchronized properly across threads. When pausing/resuming, you rely on stale cached counts instead of verifying the actual local file size, leading to incorrect range requests and overcounting.

Step 1: Fix Pause-Resume Logic Without Switching to Multiprocessing

Before jumping to multiprocessing, let’s fix your existing multi-threaded setup—it’s simpler and will resolve most issues:

1.1 Properly Terminate Connections on Pause

Each download thread should use its own requests.Session, and when pausing, you need to:

  • Use a thread-safe pause flag (like threading.Event instead of a plain boolean) for reliable signaling
  • Close all active sessions immediately to kill pending requests
  • Wait for all threads to exit gracefully before resuming

Example pause/resume code snippet:

import threading
import os

class Multidown:
    def __init__(self, url, output_path, max_connections=32):
        self.url = url
        self.output_path = output_path
        self.max_connections = max_connections
        self.pause_event = threading.Event()
        self.pause_event.set()  # Start in unpaused state
        self.download_threads = []
        self.sessions = []  # Track sessions per thread
        self.total_size = 0
        self.downloaded = 0
        self.lock = threading.Lock()
        self.mmap_obj = None

    def pause(self):
        self.pause_event.clear()
        # Close all active sessions to terminate pending requests
        for session in self.sessions:
            session.close()
        # Wait for all threads to exit
        for thread in self.download_threads:
            thread.join(timeout=5)
        # Clear thread/session lists to reset for resume
        self.download_threads.clear()
        self.sessions.clear()

    def resume(self):
        self.pause_event.set()
        # Calculate remaining bytes based on actual local file size (not cached counts)
        local_size = os.path.getsize(self.output_path) if os.path.exists(self.output_path) else 0
        self._setup_download_ranges(start_offset=local_size)
        self._start_download_threads()

1.2 Track Downloaded Bytes Safely

Use a thread-safe counter to track downloaded bytes, and enforce an upper bound to prevent progress from exceeding 100%:

def _update_downloaded(self, bytes_added):
    with self.lock:
        self.downloaded += bytes_added
        # Ensure we don't exceed total file size
        if self.downloaded > self.total_size:
            self.downloaded = self.total_size

1.3 Fix the mmap IndexError

When writing to mmap, always validate that the received data matches the expected chunk size. Trim extra bytes or log a retry if you get fewer bytes than requested:

def _download_chunk(self, session, start, end):
    headers = {'Range': f'bytes={start}-{end}'}
    try:
        response = session.get(self.url, headers=headers, stream=True)
        response.raise_for_status()
        data = b''.join(response.iter_content(chunk_size=8192))
        expected_length = end - start + 1
        
        # Trim extra bytes to avoid mmap size mismatch
        if len(data) > expected_length:
            data = data[:expected_length]
        
        # Write to mmap safely with a lock
        with self.lock:
            self.mmap_obj[start:end+1] = data
        
        self._update_downloaded(len(data))
    except Exception as e:
        # Add retry logic or error logging here
        pass

Step 2: If You Still Want to Use Multiprocessing

If you suspect thread-level connection pooling is still causing leftover bytes, switching to multiprocessing helps because each process has its own isolated socket stack. Here’s how to handle shared state:

2.1 Use Multiprocessing Shared Variables

  • For numeric values (downloaded bytes, total size), use multiprocessing.Value with lock enabled.
  • For shared lists (like download ranges), use multiprocessing.Manager().list().
  • Use multiprocessing.Event for the pause flag.

Example setup:

import multiprocessing
import requests
import os

class MultiprocessDownloader:
    def __init__(self, url, output_path, max_connections=32):
        self.url = url
        self.output_path = output_path
        self.max_connections = max_connections
        self.total_size = multiprocessing.Value('Q', 0)  # Unsigned long long for large files
        self.downloaded = multiprocessing.Value('Q', 0, lock=True)
        self.pause_event = multiprocessing.Event()
        self.pause_event.set()
        self.manager = multiprocessing.Manager()
        self.ranges = self.manager.list()

    def _download_chunk(self, start, end):
        # Create a fresh session per process with Connection: close to avoid stale connections
        session = requests.Session()
        session.headers.update({'Connection': 'close'})
        headers = {'Range': f'bytes={start}-{end}'}
        
        try:
            response = session.get(self.url, headers=headers, stream=True)
            response.raise_for_status()
            data = b''.join(response.iter_content(chunk_size=8192))
            expected_length = end - start + 1
            
            if len(data) > expected_length:
                data = data[:expected_length]
            
            # Write directly to file (avoid mmap across processes)
            with open(self.output_path, 'r+b') as f:
                f.seek(start)
                f.write(data)
            
            # Update shared downloaded count
            with self.downloaded.get_lock():
                self.downloaded.value += len(data)
                if self.downloaded.value > self.total_size.value:
                    self.downloaded.value = self.total_size.value
        finally:
            session.close()

2.2 Key Multiprocessing Notes

  • Never share requests.Session instances across processes—each process needs its own.
  • Use Connection: close to force fresh connections for every request, preventing reuse of stale sockets.
  • Avoid using mmap across processes (it’s not safe); instead, write directly to the file with seek() and write().

Step 3: Stabilize Your Test Cases

For your test1 (timed pause-resume):

  • Add a 1-2 second delay after pausing to ensure all threads/processes have exited and file writes are flushed to disk.
  • After resuming, verify the local file size matches the expected downloaded bytes before starting new downloads.
  • Run hash checks after every pause-resume cycle to confirm file integrity consistently.

Final Summary

  1. Fix the multi-threaded version first: Proper connection termination, thread-safe state tracking, and mmap bounds checking will resolve most of your issues.
  2. Switch to multiprocessing only if thread-level connection pooling is irreparably problematic: Use multiprocessing’s built-in shared variables to track state safely, and avoid shared resources like mmap.
  3. Always validate against the actual local file size when resuming—never trust cached downloaded counts.

内容的提问来源于stack exchange,提问作者Ξένη Γήινος

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 04:23:11