使用GPT-4 API是否需每次发送完整对话?与GPT-3.5-Turbo有何差异?
我是OpenAI API新手,目前使用GPT-3.5-Turbo开发,代码如下:
messages = [ {"role": "system", "content": "You’re a helpful assistant"} ] while True: content = input("User: ") if content == 'end': save_log(messages) break messages.append({"role": "user", "content": content}) completion = openai.ChatCompletion.create( model="gpt-3.5-turbo-16k", messages=messages ) chat_response = completion.choices[0].message.content print(f'ChatGPT: {chat_response}') messages.append({"role": "assistant", "content": chat_response})
测试结果:
User: who was the first person on the moon?
GPT: The first person to step foot on the moon was Neil Armstrong, an American astronaut, on July 20, 1969, as part of NASA's Apollo 11 mission.
User: how tall is he?
GPT: Neil Armstrong was approximately 5 feet 11 inches (180 cm) tall.
这种方式能实现上下文对话,但会消耗大量tokens。我听闻GPT-4与GPT-3.5-Turbo不同,可自主记忆历史消息,于是测试仅发送单条消息调用GPT-4:
completion = openai.ChatCompletion.create( model="gpt-4", messages=[{"role": "user", "content": content}] )
结果模型无法关联上下文:
User: who was the first person on the moon?
GPT: The first person on the moon was Neil Armstrong on July 20, 1969.
User: how tall is he?
GPT: Without specific context or information about who "he" refers to, I'm unable to provide an accurate answer.
我的疑问:
- 使用GPT-4 API时是否必须每次发送完整对话?
- GPT-3.5-Turbo与GPT-4在工作流程上是否存在差异?
回答
使用GPT-4 API时必须每次发送完整对话
所有OpenAI的ChatCompletion模型(包括GPT-3.5-Turbo和GPT-4)本身都不具备持久记忆能力,每次API调用都是独立请求。模型只能基于当前请求中messages参数提供的对话历史生成上下文相关回复,不存在“自主记忆历史消息”的特性。GPT-3.5-Turbo与GPT-4在工作流程上无本质差异
两者遵循完全相同的ChatCompletion API工作逻辑:依赖传入的messages数组提供对话上下文,生成回复时仅考虑该数组内的内容。差异仅体现在模型能力层面(比如GPT-4的推理能力更强、上下文窗口更大等),核心工作流程一致。
若想降低token消耗,可尝试这些优化方式:
- 对对话历史做摘要处理,仅保留关键上下文信息传入
messages - 使用大上下文窗口模型(如gpt-3.5-turbo-16k、gpt-4-32k),减少完整对话场景下的频繁截断操作
- 设计对话逻辑,将无关的历史对话从
messages中移除
内容的提问来源于stack exchange,提问作者Realchini

