C# .NET Core本地大MP4文件分块上传至外部REST API技术问询
Hey, I’ve tackled exactly this problem before—uploading large local files to external APIs via chunking, when most docs focus on web frontends. Let’s break this down into actionable steps, with a concrete Python example you can adapt to other languages if needed.
Core Concepts to Nail First
Before diving into code, make sure you align with these critical details (check your external API docs first if you haven’t):
- Chunk Size: Pick a size that balances speed and reliability (50MB-100MB is standard—avoid chunks smaller than 10MB unless the API enforces it).
- Unique File Identifier: The API needs to know which chunks belong to the same file. A file hash (MD5/SHA256) is ideal, but you could also use a custom UUID + file metadata.
- Chunk Metadata: Each upload request should include: chunk number, total number of chunks, file ID, and optionally a
Content-Rangeheader (standard for HTTP chunked uploads). - Merge Trigger: Most APIs require a final request to tell them all chunks are uploaded and ready to be assembled into the full file.
Step-by-Step Implementation (Python Example)
Python is perfect for this because it’s lightweight, has great HTTP libraries, and handles file I/O efficiently without loading the entire 4GB file into memory.
First, install the required dependency:
pip install requests
Then, here’s the implementation:
import os import hashlib import requests from typing import Optional def get_unique_file_id(file_path: str) -> str: """Generate a unique hash for the file to identify chunks in the API""" sha256_hash = hashlib.sha256() # Read file in small chunks to avoid memory overload with open(file_path, "rb") as f: for chunk in iter(lambda: f.read(4096), b""): sha256_hash.update(chunk) return sha256_hash.hexdigest() def upload_single_chunk( file_path: str, chunk_num: int, total_chunks: int, file_id: str, api_chunk_endpoint: str, chunk_size: int = 50 * 1024 * 1024 # 50MB default ) -> Optional[requests.Response]: """Upload one chunk of the file to the API""" start_byte = chunk_num * chunk_size end_byte = start_byte + chunk_size file_total_size = os.path.getsize(file_path) # Adjust end byte for the final chunk (won't fill the full chunk size) if chunk_num == total_chunks - 1: end_byte = file_total_size # Read only the current chunk from the file with open(file_path, "rb") as f: f.seek(start_byte) chunk_data = f.read(chunk_size) # Build headers the API will use to process the chunk headers = { "File-ID": file_id, "Chunk-Number": str(chunk_num + 1), # Some APIs use 1-based indexing "Total-Chunks": str(total_chunks), "Content-Range": f"bytes {start_byte}-{end_byte-1}/{file_total_size}" } try: response = requests.post( url=api_chunk_endpoint, data=chunk_data, headers=headers, timeout=30 # Adjust based on your network speed ) response.raise_for_status() # Raise error for HTTP status codes >=400 return response except requests.exceptions.RequestException as e: print(f"Failed to upload chunk {chunk_num + 1}: {str(e)}") return None def trigger_file_merge(file_id: str, api_merge_endpoint: str) -> bool: """Tell the API to assemble all chunks into the full file""" try: response = requests.post( url=api_merge_endpoint, json={"file_id": file_id} ) response.raise_for_status() print("File merged successfully!") return True except requests.exceptions.RequestException as e: print(f"Merge request failed: {str(e)}") return False def main(): # Configure your values here local_file_path = "/path/to/your/large_file.mp4" api_base_url = "https://your-external-api.com" chunk_size = 50 * 1024 * 1024 # 50MB per chunk # Calculate file metadata file_total_size = os.path.getsize(local_file_path) total_chunks = (file_total_size + chunk_size - 1) // chunk_size # Ceiling division file_id = get_unique_file_id(local_file_path) print(f"Starting upload of {local_file_path}") print(f"File ID: {file_id}, Total chunks: {total_chunks}") # Upload each chunk sequentially (add threading for parallel uploads if API allows) uploaded_chunks = [] for chunk_num in range(total_chunks): print(f"Uploading chunk {chunk_num + 1}/{total_chunks}") response = upload_single_chunk( file_path=local_file_path, chunk_num=chunk_num, total_chunks=total_chunks, file_id=file_id, api_chunk_endpoint=f"{api_base_url}/upload-chunk", chunk_size=chunk_size ) if response: uploaded_chunks.append(chunk_num) # Verify all chunks were uploaded before merging if len(uploaded_chunks) == total_chunks: trigger_file_merge( file_id=file_id, api_merge_endpoint=f"{api_base_url}/merge-file" ) else: print(f"Missing {total_chunks - len(uploaded_chunks)} chunks—retry failed ones first.") if __name__ == "__main__": main()
Key Adaptations for Your API
- Header Customization: If the API uses non-standard headers (e.g.,
X-Chunk-Indexinstead ofChunk-Number), update theheadersdictionary accordingly. - Parallel Uploads: If the API allows concurrent requests, use
concurrent.futures.ThreadPoolExecutorto upload multiple chunks at once (just be mindful of rate limits). - Breakpoint Resumption: Add logic to track uploaded chunks (e.g., save a local JSON file with
file_idand uploaded chunk numbers) so you don’t re-upload everything if the process fails halfway. - Authentication: If the API requires auth (API key, OAuth), add the required headers (e.g.,
Authorization: Bearer YOUR_TOKEN) to all requests.
内容的提问来源于stack exchange,提问作者john
相关产品推荐
相关产品推荐

