You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过OpenRouter与OpenAI库屏蔽google/gemini-2.5-flash-preview:thinking的思考步骤

问题描述

我正在使用openai Python库,通过OpenRouter API调用google/gemini-2.5-flash-preview:thinking模型进行流式聊天补全。我的目标是仅获取助手的最终回复,但该特定模型变体(:thinking)会在实际答案前,将内部思考过程(如"Thinking..."、"Examining Request..."、"Locating References...")直接插入主内容流中。

示例代码

from openai import OpenAI
import os

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ.get("OPENROUTER_API_KEY"),
)

model_name = "google/gemini-2.5-flash-preview:thinking"
messages_for_api = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain gravity briefly."}
]

try:
    response_stream = client.chat.completions.create(
        model=model_name,
        messages=messages_for_api,
        temperature=0.7,
        stream=True,
        # 注:曾尝试传入extra_body={'reasoning': {'exclude': True}},但无效
    )

    print("\nResponse Stream Chunks (Illustrative):")
    full_response = ""
    # 观察到的输出流结构示例:
    # Chunk: "Thinking..."
    # Chunk: "Okay, the user wants to know about gravity."
    # ...

    for chunk in response_stream:
         if chunk.choices and chunk.choices[0].delta and chunk.choices[0].delta.content:
            content_part = chunk.choices[0].delta.content
            print(f"Received Chunk Content: {content_part}") # 先输出思考步骤
            full_response += content_part

    print(f"\n--- 最终累积内容(包含思考步骤)--- \n{full_response}")

except Exception as e:
    print(f"\n发生错误: {e}")

已尝试的方法

  • 在系统提示中明确要求模型不要输出思考步骤,但被忽略。
  • 在create调用中传入extra_body={'reasoning': {'exclude': True}}(基于OpenRouter文档中控制"Reasoning Tokens"的说明),但对该特定模型无效。

疑问

鉴于思考步骤似乎是google/gemini-2.5-flash-preview:thinking主内容流的一部分,是否可以通过openai Python库和OpenRouter API参数,阻止该特定模型变体输出思考步骤?或者,获得干净输出的唯一可行方案是切换到不带:thinking后缀的其他模型端点(如google/gemini-1.5-flash-latest)?


解决方案

核心结论

google/gemini-2.5-flash-preview:thinking这个变体的设计初衷就是对外暴露内部思考过程,所以通过OpenRouter API参数或系统提示无法屏蔽这些内容——这是该模型变体的固有特性,而非可配置项。

可行方案

  1. 切换到不带:thinking后缀的模型
    这是最直接且可靠的方案。比如使用google/gemini-1.5-flash-latest或google/gemini-2.5-flash-preview(无后缀版本),这类模型不会输出思考步骤,直接返回最终回复。

  2. 客户端侧过滤思考内容
    如果必须使用该:thinking变体,可以在代码中对流式返回的内容进行过滤:

    • 先识别思考步骤的特征(比如以"Thinking..."、"Examining Request..."等固定前缀开头的片段)
    • 累积内容时跳过这些片段,只保留正式回复部分

    示例修改后的代码:

    from openai import OpenAI
    import os
    
    client = OpenAI(
        base_url="https://openrouter.ai/api/v1",
        api_key=os.environ.get("OPENROUTER_API_KEY"),
    )
    
    model_name = "google/gemini-2.5-flash-preview:thinking"
    messages_for_api = [
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain gravity briefly."}
    ]
    
    # 定义思考步骤的前缀列表
    thinking_prefixes = ["Thinking...", "Examining Request...", "Locating References..."]
    
    try:
        response_stream = client.chat.completions.create(
            model=model_name,
            messages=messages_for_api,
            temperature=0.7,
            stream=True,
        )
    
        print("\n过滤后的回复流:")
        full_response = ""
    
        for chunk in response_stream:
             if chunk.choices and chunk.choices[0].delta and chunk.choices[0].delta.content:
                content_part = chunk.choices[0].delta.content
                # 检查当前片段是否属于思考步骤
                is_thinking = any(content_part.startswith(prefix) for prefix in thinking_prefixes)
                if not is_thinking:
                    print(f"Received Valid Chunk: {content_part}")
                    full_response += content_part
    
        print(f"\n--- 最终干净回复 --- \n{full_response}")
    
    except Exception as e:
        print(f"\n发生错误: {e}")
    

    注意:这种方法依赖思考步骤的固定前缀,如果模型后续调整思考内容的格式,过滤逻辑可能失效,需要同步更新前缀列表。


内容的提问来源于stack exchange,提问作者MendelG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 04:52:13