You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在LangChain中为ChatGoogleGenerativeAI配置流式输出

LangChain中ChatGoogleGenerativeAI的流式输出配置方法

ChatGoogleGenerativeAI确实没有像ChatOpenAI那样在实例初始化时提供streaming参数,但可以通过以下两种方式实现逐token的流式响应:

1. 直接调用stream()方法

无需在实例创建阶段做额外配置,直接调用stream()方法即可获取逐token输出:

from langchain.chat_models import ChatGoogleGenerativeAI
from langchain.schema import HumanMessage
from langchain_google_genai import HarmCategory, HarmBlockThreshold

# 创建ChatGoogleGenerativeAI实例
llm = ChatGoogleGenerativeAI(
    model="gemini-pro",
    safety_settings={
        HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT: HarmBlockThreshold.BLOCK_NONE,
    },
)

# 调用stream方法,遍历获取每个token块
for chunk in llm.stream([HumanMessage(content="请简述LangChain的主要用途")]):
    print(chunk.content, end="", flush=True)

2. 结合回调函数处理token输出

如果需要自定义每个token的处理逻辑(比如实时渲染到界面),可以创建回调类并在调用stream()时传入:

from langchain.chat_models import ChatGoogleGenerativeAI
from langchain.schema import HumanMessage
from langchain.callbacks.base import BaseCallbackHandler
from langchain_google_genai import HarmCategory, HarmBlockThreshold
import sys

# 自定义流式回调类
class StreamingTokenHandler(BaseCallbackHandler):
    def on_llm_new_token(self, token: str, **kwargs) -> None:
        # 这里可以添加自定义逻辑,比如写入前端、保存到文件等
        sys.stdout.write(token)
        sys.stdout.flush()

# 创建实例
llm = ChatGoogleGenerativeAI(
    model="gemini-pro",
    safety_settings={
        HarmCategory.HARM_CATEGORY_DANGEROUS_CONTENT: HarmBlockThreshold.BLOCK_NONE,
    },
)

# 调用stream时传入回调
llm.stream(
    [HumanMessage(content="请详细介绍Gemini模型的特点")],
    callbacks=[StreamingTokenHandler()]
)

关键说明

ChatGoogleGenerativeAI的设计逻辑与ChatOpenAI不同:它没有将流式开关放在实例初始化阶段,而是通过stream()方法显式触发流式输出。上述两种方式都能实现逐token响应,区别在于是否需要自定义token的处理逻辑。

内容的提问来源于stack exchange,提问作者Mohamed Anser Ali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 19:32:12