You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Langchain4j中如何复用嵌入文档实现LLM少样本查询?

解决方案:通过系统提示复用共同文档上下文减少Token消耗

可以通过将共享的文档上下文移至**系统提示词(System Message)**中,让LLM将其作为所有少样本示例和实际查询的共同参考,彻底避免重复嵌入文档内容导致的Token浪费。具体实现如下:

核心优化思路

把原本重复放在每个用户消息里的文档信息,统一放到系统消息中,作为模型处理所有后续消息的全局上下文。少样本示例只需保留「问题-预期回复」的结构,实际查询也仅传递问题,无需重复携带文档内容。

优化后的Java LangChain代码实现

// 1. 加载并处理文档(原有逻辑不变)
Document document = loadDocument(toPath("file:///filepath/filename.pdf"));
DocumentByRegexSplitter splitter = new DocumentByRegexSplitter(regex, joiner, maxCharLimit, maxOverlap, subSplitter);
// ... 文档分割、嵌入、获取相关向量的逻辑保持不变
String information = relevantEmbeddings.stream()
        .map(match -> match.embedded().text())
        .collect(joining("\n\n"));

// 2. 重构提示词:将文档信息移至系统提示,用户提示仅保留问题部分
// 系统提示:明确任务要求 + 共享文档上下文
SystemMessage systemMessage = new SystemMessage(
        "你需要基于以下提供的文档信息,尽可能准确地回答问题。\n"
        + "文档信息:\n"
        + information
);

// 用户提示模板:仅包含问题,无需重复文档内容
PromptTemplate userPromptTemplate = PromptTemplate.from(
        "Question:\n{{question}}"
);

// 3. 构造聊天消息列表
List<ChatMessage> chatMessages = new ArrayList<>();
chatMessages.add(systemMessage);

// 添加少样本示例:仅传递问题和预期回复
Map<String, Object> trainingVars = new HashMap<>();
trainingVars.put("question", trainingQuestion);
Prompt trainingPrompt = userPromptTemplate.apply(trainingVars);
chatMessages.add(trainingPrompt.toUserMessage());
chatMessages.add(new AiMessage("Expected Response"));

// 添加实际查询:仅传递问题
Map<String, Object> actualVars = new HashMap<>();
actualVars.put("question", actualQuestion);
Prompt actualPrompt = userPromptTemplate.apply(actualVars);
chatMessages.add(actualPrompt.toUserMessage());

// 4. 生成响应
AiMessage response = chatModel.generate(chatMessages);

为什么这能解决Token问题?

  • 系统提示会被LLM视为全局上下文,所有后续的少样本示例和实际查询都会基于这个上下文处理,无需重复嵌入文档内容。
  • 原本每个用户消息都携带的大段文档信息,现在只需要传递一次,Token消耗直接减少(少样本数量-1)倍的文档内容Token数。

额外优化建议

如果文档内容本身仍较长,可以进一步:

  • 优化向量检索逻辑,仅保留与少样本示例、实际查询高度相关的文档片段,而非全部相关内容。
  • 对文档片段进行二次精简(比如用LLM提取关键信息),进一步压缩上下文长度。

内容的提问来源于stack exchange,提问作者Sachu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 00:32:32