You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google Cloud Run上终止Django的StreamingHttpResponse?

Google Cloud Run 上 Django 流式响应的早期终止方案

在现有环境下完全可以实现,以下是两种无需额外依赖的可行方案:

方案一:利用客户端断开的信号被动终止

  • 当用户主动关闭页面或终止请求时,Google Cloud Run会向容器发送连接断开的信号,Django的请求对象会触发BrokenPipeError异常,也可通过请求头判断连接状态。
  • 在流式生成器中捕获异常或检查请求状态,一旦检测到连接断开,立即停止调用GPT API并终止流。
  • 示例代码:
import sys
from django.http import StreamingHttpResponse

def stream_gpt_response(request):
    def generator():
        try:
            # 调用GPT的流式接口获取数据块
            for chunk in call_gpt_streaming_api():
                # 检查连接是否已关闭
                if request.META.get('HTTP_CONNECTION') == 'close':
                    break
                yield chunk
                # 强制刷新缓冲区,确保信号及时传递
                sys.stdout.flush()
        except BrokenPipeError:
            # 客户端断开,直接终止生成器
            pass
    return StreamingHttpResponse(generator(), content_type='text/plain')
  • 注意点:需合理设置Cloud Run的请求超时时间,避免容器提前被回收;生成器内的检查频率要适中,平衡性能与终止延迟。

方案二:基于本地内存缓存的主动终止

  • 给每个流式请求生成唯一request_id,存入Django的本地内存缓存(LocMemCache),初始标记为active。
  • 提供一个独立的终止接口,用户传入request_id后,将缓存中的标记改为terminated。
  • 在流式生成器循环中,每次返回数据前检查request_id的状态,若标记为终止则停止流。
  • 示例代码:
import uuid
from django.core.cache import cache
from django.http import StreamingHttpResponse, HttpResponse

def stream_gpt_response(request):
    request_id = str(uuid.uuid4())
    # 缓存标记设置5分钟超时,避免内存泄漏
    cache.set(request_id, 'active', timeout=300)
    
    def generator():
        try:
            for chunk in call_gpt_streaming_api():
                if cache.get(request_id) != 'active':
                    cache.delete(request_id)
                    break
                yield chunk
        finally:
            # 无论正常结束还是终止,都清理缓存
            cache.delete(request_id)
    
    response = StreamingHttpResponse(generator(), content_type='text/plain')
    # 将request_id通过响应头返回给客户端,用于后续终止请求
    response['X-Request-ID'] = request_id
    return response

def terminate_stream(request):
    request_id = request.GET.get('request_id')
    if request_id:
        cache.set(request_id, 'terminated', timeout=60)
    return HttpResponse(status=204)
  • 注意点:LocMemCache是进程内缓存,若Cloud Run有多个实例,终止请求需要发送到对应实例。可以让客户端记录响应头中的实例标识(如X-Cloud-Run-Instance),终止时指定实例,适合低并发或实例数量较少的场景。

这两种方案都不需要WebSocket、Redis或VPC相关服务,完全基于Cloud Run的Django容器本身实现,满足你的成本要求。

内容的提问来源于stack exchange,提问作者M S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 15:41:07