You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Django StreamingHttpResponse返回S3大文件无流数据问题排查

问题分析:S3流式返回PDF时请求Pending无分块数据

问题场景

现有一段从S3获取PDF并流式返回给前端的Django代码:

def stream_pdf_from_s3(request, file_key):
    s3_client = boto3.client('s3')

    try:
        response = s3_client.get_object(Bucket=settings.AWS_STORAGE_BUCKET_NAME, Key=file_key)
        pdf_stream = response['Body']

        # Use iter_chunks() for efficient streaming
        return StreamingHttpResponse(pdf_stream.iter_chunks(chunk_size=65536), content_type='application/pdf')

    except Exception as e:
        return HttpResponse(f"Error fetching PDF: {e}", status=500)

注:原代码中chunk_size的写法65,536存在语法错误,已修正为65536

遇到的问题:浏览器网络面板中请求长期处于pending状态,无分块字节返回,未实现立即返回分块数据的预期效果。

问题原因

  • boto3的iter_chunks并非真正流式拉取:get_object默认会先将整个S3文件下载到本地缓存,再通过iter_chunks返回数据块。这意味着文件未完全下载前,前端不会收到任何响应,导致请求一直pending。
  • 缺少分块传输响应头:Django的StreamingHttpResponse默认不会自动添加Transfer-Encoding: chunked头,部分浏览器或代理需要该头才能识别分块传输并逐步接收数据。
  • 缓存中间件阻断流式逻辑:若项目配置了缓存类中间件(如UpdateCacheMiddleware),这类中间件会等待整个响应完成后再发送给客户端,直接破坏流式传输的流程。

修正方案

1. 用S3范围请求实现真正流式拉取

通过范围请求每次拉取固定大小的块,避免一次性下载整个文件:

from django.http import StreamingHttpResponse, HttpResponse
import boto3
from django.conf import settings
from django.views.decorators.cache import never_cache

@never_cache
def stream_pdf_from_s3(request, file_key):
    s3_client = boto3.client('s3')
    chunk_size = 65536  # 64KB

    try:
        # 获取文件总大小
        head_response = s3_client.head_object(Bucket=settings.AWS_STORAGE_BUCKET_NAME, Key=file_key)
        file_size = head_response['ContentLength']

        def s3_stream_generator():
            start = 0
            while start < file_size:
                end = min(start + chunk_size - 1, file_size - 1)
                range_response = s3_client.get_object(
                    Bucket=settings.AWS_STORAGE_BUCKET_NAME,
                    Key=file_key,
                    Range=f'bytes={start}-{end}'
                )
                yield range_response['Body'].read()
                start = end + 1

        response = StreamingHttpResponse(s3_stream_generator(), content_type='application/pdf')
        # 添加分块传输标识头
        response['Transfer-Encoding'] = 'chunked'
        # 可选:添加文件大小头,让浏览器显示下载进度
        response['Content-Length'] = str(file_size)
        return response

    except Exception as e:
        return HttpResponse(f"Error fetching PDF: {e}", status=500)

2. 配置生产环境WSGI服务器

如果使用生产级WSGI服务器,需确保启用流式支持:

  • Gunicorn:添加--sendfile off参数,禁用内核级sendfile缓存
  • uWSGI:配置wsgi-disable-file-wrapper = true,关闭文件包装器缓存

内容的提问来源于stack exchange,提问作者Azima

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 09:24:52