如何使用Requests库以低内存占用方式分块上传大文件?
Great question! I’ve dealt with exactly this kind of large-file upload constraint before—balancing API limits with memory efficiency is tricky, but there’s a clean way to do it with Requests that mirrors the httplib approach you mentioned.
The key insight here is that Requests supports iterable objects as the data parameter. Instead of loading an entire 10MB chunk into memory, we can create a generator that reads small bytes from the file (4096/8192 at a time) and yields them until we’ve sent exactly 10MB. Requests will iterate over this generator and send the data incrementally, without buffering the full 10MB in memory.
Step-by-Step Solution
First, we’ll build a generator function that controls how much data we send per request. This generator will read small chunks from the file, stop once we’ve hit the 10MB target, and let Requests handle sending each tiny piece as it’s read.
Here’s a complete working example:
import os import requests def stream_10mb_chunk(file_obj, small_chunk_size=4096): """Generator that yields small chunks from a file until 10MB is reached.""" target_total = 10 * 1024 * 1024 # 10MB in bytes bytes_sent = 0 while bytes_sent < target_total: # Calculate how much we can read in this iteration (don't exceed remaining target) read_size = min(small_chunk_size, target_total - bytes_sent) chunk = file_obj.read(read_size) if not chunk: break # End of file reached early (handle this based on your API's rules) bytes_sent += len(chunk) yield chunk # Usage example with open("your_5gb_file.bin", "rb") as large_file: total_file_size = os.path.getsize("your_5gb_file.bin") full_10mb_chunks = total_file_size // (10 * 1024 * 1024) # Process all full 10MB chunks for _ in range(full_10mb_chunks): # Create our streaming generator for this request chunk_stream = stream_10mb_chunk(large_file) # Send the request with the stream, and explicitly set Content-Length response = requests.post( "https://your-api-endpoint.com/upload", data=chunk_stream, headers={"Content-Length": str(10 * 1024 * 1024)} ) response.raise_for_status() # Handle any HTTP errors here print(f"Completed chunk {_+1}/{full_10mb_chunks}") # Handle remaining data (if any) remaining_bytes = total_file_size % (10 * 1024 * 1024) if remaining_bytes > 0: # Note: Your API requires a minimum 10MB per request, so you may need to pad with zeros here # Example padding logic: final_data = large_file.read(remaining_bytes) final_data += b"\x00" * (10 * 1024 * 1024 - remaining_bytes) response = requests.post( "https://your-api-endpoint.com/upload", data=final_data, headers={"Content-Length": str(10 * 1024 * 1024)} ) response.raise_for_status() print("Completed final padded chunk")
How This Works
- Memory Efficiency: The generator only reads 4096 bytes at a time from the file, so your memory usage stays low (just a few KB per request, not 10MB).
- API Compliance: Each request sends exactly 10MB (either full file chunks or padded remaining data), which meets your API’s minimum/maximum requirements.
- Requests Compatibility: Requests natively handles iterable
datainputs—under the hood, it’s doing something very similar tohttplib’ssend()loop, but wrapped in Requests’ easier-to-use interface.
Key Notes
- Content-Length Header: It’s critical to explicitly set this header. Without it, Requests will use chunked transfer encoding, which some APIs don’t support (and your API expects a fixed 10MB payload per request).
- Final Chunk Handling: If your file isn’t a perfect multiple of 10MB, you’ll need to pad the last chunk to reach 10MB (as shown) to comply with the API’s minimum size rule.
- Error Handling: Add retries or logging as needed—network issues can happen with large uploads, so you’ll want to handle failed requests gracefully.
内容的提问来源于stack exchange,提问作者Simplecode

