如何通过OpenRouter与OpenAI库屏蔽google/gemini-2.5-flash-preview:thinking的思考步骤
我正在使用openai Python库,通过OpenRouter API调用google/gemini-2.5-flash-preview:thinking模型进行流式聊天补全。我的目标是仅获取助手的最终回复,但该特定模型变体(:thinking)会在实际答案前,将内部思考过程(如"Thinking..."、"Examining Request..."、"Locating References...")直接插入主内容流中。
示例代码
from openai import OpenAI import os client = OpenAI( base_url="https://openrouter.ai/api/v1", api_key=os.environ.get("OPENROUTER_API_KEY"), ) model_name = "google/gemini-2.5-flash-preview:thinking" messages_for_api = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Explain gravity briefly."} ] try: response_stream = client.chat.completions.create( model=model_name, messages=messages_for_api, temperature=0.7, stream=True, # 注:曾尝试传入extra_body={'reasoning': {'exclude': True}},但无效 ) print("\nResponse Stream Chunks (Illustrative):") full_response = "" # 观察到的输出流结构示例: # Chunk: "Thinking..." # Chunk: "Okay, the user wants to know about gravity." # ... for chunk in response_stream: if chunk.choices and chunk.choices[0].delta and chunk.choices[0].delta.content: content_part = chunk.choices[0].delta.content print(f"Received Chunk Content: {content_part}") # 先输出思考步骤 full_response += content_part print(f"\n--- 最终累积内容(包含思考步骤)--- \n{full_response}") except Exception as e: print(f"\n发生错误: {e}")
已尝试的方法
- 在系统提示中明确要求模型不要输出思考步骤,但被忽略。
- 在
create调用中传入extra_body={'reasoning': {'exclude': True}}(基于OpenRouter文档中控制"Reasoning Tokens"的说明),但对该特定模型无效。
疑问
鉴于思考步骤似乎是google/gemini-2.5-flash-preview:thinking主内容流的一部分,是否可以通过openai Python库和OpenRouter API参数,阻止该特定模型变体输出思考步骤?或者,获得干净输出的唯一可行方案是切换到不带:thinking后缀的其他模型端点(如google/gemini-1.5-flash-latest)?
核心结论
google/gemini-2.5-flash-preview:thinking这个变体的设计初衷就是对外暴露内部思考过程,所以通过OpenRouter API参数或系统提示无法屏蔽这些内容——这是该模型变体的固有特性,而非可配置项。
可行方案
切换到不带
:thinking后缀的模型
这是最直接且可靠的方案。比如使用google/gemini-1.5-flash-latest或google/gemini-2.5-flash-preview(无后缀版本),这类模型不会输出思考步骤,直接返回最终回复。客户端侧过滤思考内容
如果必须使用该:thinking变体,可以在代码中对流式返回的内容进行过滤:- 先识别思考步骤的特征(比如以"Thinking..."、"Examining Request..."等固定前缀开头的片段)
- 累积内容时跳过这些片段,只保留正式回复部分
示例修改后的代码:
from openai import OpenAI import os client = OpenAI( base_url="https://openrouter.ai/api/v1", api_key=os.environ.get("OPENROUTER_API_KEY"), ) model_name = "google/gemini-2.5-flash-preview:thinking" messages_for_api = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Explain gravity briefly."} ] # 定义思考步骤的前缀列表 thinking_prefixes = ["Thinking...", "Examining Request...", "Locating References..."] try: response_stream = client.chat.completions.create( model=model_name, messages=messages_for_api, temperature=0.7, stream=True, ) print("\n过滤后的回复流:") full_response = "" for chunk in response_stream: if chunk.choices and chunk.choices[0].delta and chunk.choices[0].delta.content: content_part = chunk.choices[0].delta.content # 检查当前片段是否属于思考步骤 is_thinking = any(content_part.startswith(prefix) for prefix in thinking_prefixes) if not is_thinking: print(f"Received Valid Chunk: {content_part}") full_response += content_part print(f"\n--- 最终干净回复 --- \n{full_response}") except Exception as e: print(f"\n发生错误: {e}")注意:这种方法依赖思考步骤的固定前缀,如果模型后续调整思考内容的格式,过滤逻辑可能失效,需要同步更新前缀列表。
内容的提问来源于stack exchange,提问作者MendelG

