You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GAE部署FastAPI:StreamingResponse未流式返回VertexAI响应问题

问题:GAE部署FastAPI流式接口无法返回VertexAI流式响应,而是一次性返回完整内容

用户在Google App Engine(GAE)平台部署了调用VertexAI流式生成的FastAPI接口,期望流式获取响应,但实际调用时会一次性收到完整内容。相关代码如下:

生成响应的prompt_ai函数

import vertexai
import os
import time
from vertexai.language_models import TextGenerationModel

def prompt_ai(prompt):
    vertexai.init(project="XXX-YYYY", location="ZZ-PPPP")
    parameters = {
        "max_output_tokens": 1024,
        "temperature": 0.2,
        "top_p": 0.8,
        "top_k": 40
    }
    model = TextGenerationModel.from_pretrained("text-bison")
    responses = model.predict_streaming(
        prompt,
        **parameters
    )
    results = []
    #print ("===========>>>> GETTING VERTEX RESPONSE <<<<================")
    for response in responses:
        text_chunk = str(response)
        yield text_chunk

FastAPI端点代码

async def search(ai_prompt: str): 
  return StreamingResponse(prompt_ai(ai_prompt), media_type='text/event-stream')

本地调用脚本

import requests

url = "https://myGCPdomain.appspot.com/search"
params = {
    "ai_prompt": "Tell me something funny",
}

headers = {
    "Authorization": "Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6Ietc..."
}

response = requests.post(url, params=params, headers=headers, stream=True)

for chunk in response.iter_lines():
    if chunk:
        print(chunk.decode("utf-8"))

问题原因

  • GAE默认响应缓冲:GAE标准环境的前端服务器会默认缓冲所有响应,直到响应完全生成后才发送给客户端,即使使用StreamingResponse也会被这个机制拦截,导致流式失效。
  • 缺少禁用缓冲的响应头:未在响应中添加明确禁用缓冲的头部字段,GAE无法识别需要流式传输。
  • 同步生成器的兼容性问题:prompt_ai是同步生成器函数,FastAPI在GAE环境下处理同步生成器时,可能因同步阻塞触发整体缓冲,无法逐块发送内容。

解决方案

1. 添加禁用缓冲的响应头

在StreamingResponse中添加Cache-Control和X-Accel-Buffering头,强制GAE关闭缓冲:

async def search(ai_prompt: str):
    headers = {
        "Cache-Control": "no-cache",
        "X-Accel-Buffering": "no"
    }
    return StreamingResponse(prompt_ai(ai_prompt), media_type='text/event-stream', headers=headers)

2. 将同步生成器转为异步生成器

把同步的生成逻辑包装为异步生成器,避免GAE因同步操作触发缓冲,同时可添加微小延迟确保块被及时发送:

import asyncio
import vertexai
from vertexai.language_models import TextGenerationModel

async def prompt_ai_async(prompt):
    vertexai.init(project="XXX-YYYY", location="ZZ-PPPP")
    parameters = {
        "max_output_tokens": 1024,
        "temperature": 0.2,
        "top_p": 0.8,
        "top_k": 40
    }
    model = TextGenerationModel.from_pretrained("text-bison")
    responses = model.predict_streaming(prompt, **parameters)
    
    async def generate_chunks():
        for response in responses:
            yield str(response)
            # 添加微小延迟,避免GAE合并过小的响应块
            await asyncio.sleep(0.01)
    
    return generate_chunks()

# 修改FastAPI端点
async def search(ai_prompt: str):
    headers = {
        "Cache-Control": "no-cache",
        "X-Accel-Buffering": "no"
    }
    return StreamingResponse(prompt_ai_async(ai_prompt), media_type='text/event-stream', headers=headers)

3. 检查GAE环境配置

  • 若使用GAE标准环境,确保app.yaml中没有开启强制缓冲的配置项;
  • 若使用GAE灵活环境,需确认配套的nginx配置(如有)已禁用响应缓冲。

内容的提问来源于stack exchange,提问作者APIS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 00:00:58