You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从LangChain内存中移除条目?解决上下文长度超限报错问题

解决LangChain内存上下文超出模型Token限制的方案

一、移除内存中的最后一个对话对象

LangChain的内存类(如ConversationBufferMemory)底层通常通过chat_memory.messages列表存储对话内容,直接操作这个列表就能移除最后一条消息:

# 假设memory是已初始化的ConversationBufferMemory实例
if memory.chat_memory.messages:
    memory.chat_memory.messages.pop()  # 移除最后一条对话记录

其他基于ChatMessageHistory的内存实现(比如ConversationSummaryMemory)也可以用同样的方式操作,只要能访问到存储对话的列表即可。

二、应用滑动窗口策略

1. 使用内置的ConversationBufferWindowMemory

LangChain原生提供了滑动窗口内存实现,通过k参数指定保留最近的k轮对话,自动丢弃更早的记录:

from langchain.memory import ConversationBufferWindowMemory

# 保留最近3轮对话,超过的自动删除
memory = ConversationBufferWindowMemory(k=3)

2. 基于Token数动态调整窗口(更精准)

如果需要根据实际Token用量动态调整,而不是固定轮数,可以自定义修剪逻辑。借助tiktoken库计算每条消息的Token数,从最早的消息开始移除,直到总Token数符合模型限制:

from langchain.memory import ConversationBufferMemory
import tiktoken

def trim_memory_to_token_limit(memory, model="gpt-3.5-turbo", max_total_tokens=4047):
    # 初始化对应模型的Token编码器
    encoder = tiktoken.encoding_for_model(model)
    total_tokens = 0
    retained_messages = []

    # 从最新消息往前累加Token,直到达到限制
    for msg in reversed(memory.chat_memory.messages):
        msg_token_count = len(encoder.encode(msg.content))
        if total_tokens + msg_token_count > max_total_tokens:
            break
        retained_messages.append(msg)
        total_tokens += msg_token_count

    # 反转列表恢复原有顺序,替换原内存中的消息
    memory.chat_memory.messages = list(reversed(retained_messages))

# 使用示例
memory = ConversationBufferMemory()
# 假设内存中已积累大量对话,调用函数修剪
trim_memory_to_token_limit(memory)

注意:如果对话中包含函数调用,需要提前预留出函数对应的Token量(比如你报错里的50个Token),所以max_total_tokens要设置为模型最大限制减去函数Token数。

内容的提问来源于stack exchange,提问作者Areza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 23:53:21