You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flask API调用本地LLM流式响应问题:stdout正常yield失效

Flask流式返回LLM响应:解决yield在LangChain回调中无效的问题

问题根源分析

  • stdout.write(token)生效的原因:它直接往标准输出写入内容,Flask开发服务器会临时将标准输出内容作为响应返回,但这是非常规实现,生产环境(如Gunicorn部署)不会生效,且无法控制响应格式。
  • yield token无效的原因:LangChain的BaseCallbackHandler的on_llm_new_token是事件通知接口,仅用于接收新token的触发信号,而非返回生成器。你在该方法中yield的内容不会被任何代码迭代,无法传递给Flask的Response对象;同时chain.run()是同步执行方法,会等待LLM生成完所有内容才返回,完全未利用流式特性。

解决方案

方案一:直接使用LangChain的stream方法(推荐)

LangChain的LLMChain提供stream方法,直接返回包含token的生成器,无需依赖回调。将生成器传入Flask的Response即可实现流式输出。

修改后的Flask代码:

from langchain.llms import CTransformers
from langchain.chains import LLMChain
from flask import Flask, Response
from langchain.prompts import PromptTemplate

app = Flask(__name__)

@app.get('/<query>')
def stream_response(query):
    # 初始化LLM,开启流式模式
    model_id = 'TheBloke/Mistral-7B-codealpaca-lora-GGUF'
    config = {'temperature': 0.00, 'context_length': 4000} 
    llm = CTransformers(
        model=model_id, 
        model_type='mistral',
        config=config,
        stream=True
    )

    # 构建Prompt和Chain
    prompt = PromptTemplate.from_template("you are an assistant answer the following : {query}")
    chain = LLMChain(llm=llm, prompt=prompt)

    # 定义生成器函数,逐个返回token
    def generate_tokens():
        for chunk in chain.stream({"query": query}):
            yield chunk['text']
    
    # 返回流式响应,指定文本内容类型
    return Response(generate_tokens(), content_type='text/plain')
   
if __name__ == '__main__':
   app.run(port=8000, threaded=True)

对应的客户端调用代码(需迭代iter_content而非直接取resp.text):

import requests

# 替换为实际查询内容
with requests.get('http://127.0.0.1:8000/Explain what Flask is in simple terms', stream=True) as resp:
    # 逐个接收流式chunk并实时打印
    for chunk in resp.iter_content(chunk_size=None, decode_unicode=True):
        if chunk:
            print(chunk, end='', flush=True)

方案二:回调+队列实现(适合必须使用回调的场景)

如果需要在回调中做额外处理(如日志、token过滤),可以用队列存储token,在视图函数中通过生成器从队列读取并返回。

修改后的代码:

from langchain.llms import CTransformers
from langchain.callbacks.base import BaseCallbackHandler
from langchain.chains import LLMChain
from flask import Flask, Response
from langchain.prompts import PromptTemplate
from queue import Queue
import threading
from typing import Any

# 自定义回调,将token存入队列
class QueueCallbackHandler(BaseCallbackHandler):
    def __init__(self, queue: Queue):
        self.queue = queue

    def on_llm_new_token(self, token: str, **kwargs: Any) -> None:
        self.queue.put(token)

    def on_llm_end(self, *args, **kwargs) -> None:
        # 发送结束信号,通知生成器停止
        self.queue.put(None)

app = Flask(__name__)

@app.get('/<query>')
def stream_response(query):
    token_queue = Queue()
    handler = QueueCallbackHandler(token_queue)
    
    # 初始化LLM和Chain
    model_id = 'TheBloke/Mistral-7B-codealpaca-lora-GGUF'
    config = {'temperature': 0.00, 'context_length': 4000} 
    llm = CTransformers(
        model=model_id, 
        model_type='mistral',
        config=config,
        stream=True,
        callbacks=[handler]
    )

    prompt = PromptTemplate.from_template("you are an assistant answer the following : {query}")
    chain = LLMChain(llm=llm, prompt=prompt)

    # 启动线程执行chain.run,避免阻塞生成器
    threading.Thread(target=lambda: chain.run(query)).start()

    # 生成器从队列读取token并返回
    def generate_tokens():
        while True:
            token = token_queue.get()
            if token is None:
                break
            yield token
    
    return Response(generate_tokens(), content_type='text/plain')
   
if __name__ == '__main__':
   app.run(port=8000, threaded=True)

内容的提问来源于stack exchange,提问作者Joy Maitra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 07:53:17