如何用Python-Flask实现异步大文件上传以解决超时问题?
Hey there! Let's work through this Flask async upload problem you're facing with those long audio files—totally get why the default 300-second timeout is causing headaches. Here's a practical, step-by-step guide to fix this properly:
First off, make sure you're running Flask 2.0 or newer—it has built-in async route support, which lets you return a response immediately while handling the heavy lifting in the background. Here's a quick example:
from flask import Flask, request import asyncio app = Flask(__name__) @app.route('/upload-audio', methods=['POST']) async def upload_audio(): # Grab the uploaded file first audio_file = request.files.get('audio') if not audio_file: return {"error": "No audio file provided"}, 400 # Offload the processing to a background async task asyncio.create_task(process_audio_async(audio_file)) # Return an immediate response so the client doesn't hit timeout return {"status": "Audio upload received—processing started!"}, 202 async def process_audio_async(audio_file): # This is where your actual audio processing logic goes # Example: Save the file to disk first save_path = f"/path/to/your/storage/{audio_file.filename}" audio_file.save(save_path) # Simulate a long-running process (replace with your actual code) await asyncio.sleep(300) # 5 minutes of processing # Run your audio processing commands here—like using `ffmpeg` or a custom script # subprocess.run(["ffmpeg", "-i", save_path, f"{save_path}_processed.mp3"], check=True) print(f"Done processing {audio_file.filename}!")
Key benefit: The client gets a 202 Accepted response right away, so they never hit the HTTP timeout. The processing happens quietly in the background.
If your audio processing takes really long (30+ minutes) or you need fault tolerance (e.g., if your Flask server restarts mid-processing), a task queue like Celery paired with Redis/RabbitMQ is the way to go. Asyncio background tasks can get lost if the server restarts—Celery persists tasks so they don't vanish.
Step 1: Install Dependencies
pip install celery redis
Step 2: Configure Celery and Flask
# celery_config.py from celery import Celery def make_celery(app): celery = Celery( app.import_name, broker=app.config['CELERY_BROKER_URL'], backend=app.config['CELERY_RESULT_BACKEND'] ) celery.conf.update(app.config) return celery # app.py from flask import Flask, request from celery_config import make_celery import subprocess import os app = Flask(__name__) # Configure Celery to use Redis as broker/backend app.config.update( CELERY_BROKER_URL='redis://localhost:6379/0', CELERY_RESULT_BACKEND='redis://localhost:6379/0' ) celery = make_celery(app) @app.route('/upload-audio', methods=['POST']) def upload_audio(): audio_file = request.files.get('audio') if not audio_file: return {"error": "No audio file provided"}, 400 # Save the file to a temporary location (Celery can't pass file objects directly) temp_path = f"/tmp/{audio_file.filename}" audio_file.save(temp_path) # Send the processing task to the Celery queue task = process_audio.delay(temp_path, audio_file.filename) # Return task ID so the client can check status later return { "status": "Processing started", "task_id": task.id, "tip": "Check progress at /task/<your-task-id>" }, 202 @celery.task(bind=True) def process_audio(self, temp_path, filename): # Your audio processing logic here try: # Example: Convert audio format with ffmpeg output_path = f"/path/to/processed/{filename}_converted.mp3" subprocess.run( ["ffmpeg", "-i", temp_path, output_path], check=True, stdout=subprocess.PIPE, stderr=subprocess.PIPE ) # Clean up temporary file os.remove(temp_path) return {"status": "completed", "processed_file": output_path} except subprocess.CalledProcessError as e: # Handle errors gracefully os.remove(temp_path) self.update_state(state='FAILURE', info=str(e.stderr)) raise @app.route('/task/<task_id>') def get_task_status(task_id): task = process_audio.AsyncResult(task_id) if task.state == 'PENDING': return {"state": "pending", "status": "Still processing your audio..."} elif task.state != 'FAILURE': return {"state": task.state, "result": task.result} else: return {"state": "failed", "error": str(task.info)}
Step 3: Run Celery Worker
In a separate terminal, start the Celery worker to process tasks:
celery -A app.celery worker --loglevel=info
Why this works: Tasks are stored in Redis, so even if your Flask server restarts, the worker will pick up where it left off. Clients can also query the task status to see when processing is done.
Even with async handling, if the file upload itself takes longer than 300 seconds (e.g., large audio over slow internet), you need to adjust timeouts at every layer:
- Flask Dev Server: Run with
flask run --timeout 600(sets 10-minute timeout) - Gunicorn: Start with
gunicorn --timeout 600 app:app - Nginx (if using reverse proxy): Add these settings to your server block:
server { # ... other config ... client_max_body_size 100M; # Adjust based on your max audio file size location / { proxy_pass http://localhost:8000; proxy_read_timeout 600s; # 10-minute timeout for proxy responses proxy_connect_timeout 600s; proxy_send_timeout 600s; } }
For extra large files (1GB+), consider implementing shard uploads on the client side. Split the audio into small chunks (e.g., 10MB each), upload them one by one, and then have your Flask server merge them into a single file once all chunks are received. This avoids single-request timeouts entirely and adds support for resuming interrupted uploads.
Pick the approach that fits your needs: async routes are perfect for quick, low-risk setups, while Celery is ideal for production where reliability matters most. Don't forget to test with your actual audio file sizes to make sure all timeouts are tuned correctly!
内容的提问来源于stack exchange,提问作者Jubaer Hossain

