You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Haystack 2.X中将图片作为用户提示词加入聊天机器人流水线?

在Haystack流水线中实现图片提问功能

你之前的方案不生效,核心原因是PromptBuilder仅用于构建纯文本提示词,无法处理多模态内容;同时需要使用OpenAI的多模态模型(如gpt-4o、gpt-4-vision-preview)才能解析图片输入。以下是正确的实现方式:

核心思路

  1. 切换到支持多模态的OpenAI模型,放弃纯文本模型
  2. 绕过PromptBuilder,直接向OpenAIGenerator传入符合OpenAI多模态API格式的结构化消息
  3. 将图片转换为base64编码(或提供可公开访问的URL),嵌入到请求内容中

完整代码示例

from haystack import Pipeline
from haystack.utils import Secret
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever
from haystack.components.generators import OpenAIGenerator
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack import Document
import base64

# 初始化文档存储与检索器
docstore = InMemoryDocumentStore()
docstore.write_documents([
    Document(content="Rome is the capital of Italy"), 
    Document(content="Paris is the capital of France")
])

# 工具函数:将本地图片转为base64编码
def image_to_base64(image_path):
    with open(image_path, "rb") as image_file:
        return base64.b64encode(image_file.read()).decode("utf-8")

# 替换为你的图片路径与问题
image_base64 = image_to_base64("your_image.jpg")
query = "描述这张图片的内容,同时告诉我法国的首都是什么?"

# 初始化支持多模态的LLM
llm = OpenAIGenerator(
    api_key=Secret.from_token("YOUR_OPENAI_API_KEY"),
    model="gpt-4o"  # 或gpt-4-vision-preview
)

# 构建流水线
pipe = Pipeline()
pipe.add_component("retriever", InMemoryBM25Retriever(document_store=docstore))
pipe.add_component("llm", llm)

# 构造多模态请求内容
def build_multimodal_input(retrieved_docs, query, image_base64=None):
    # 拼接检索到的上下文文本
    context = "\n".join([doc.content for doc in retrieved_docs])
    
    # 构建OpenAI标准格式的消息内容
    message_content = [
        {"type": "text", "text": f"参考以下上下文回答问题:{context}\n\n问题:{query}"}
    ]
    
    # 加入图片内容(如果有)
    if image_base64:
        message_content.append({
            "type": "image_url",
            "image_url": {"url": f"data:image/jpeg;base64,{image_base64}"}
        })
    
    return {"messages": [{"role": "user", "content": message_content}]}

# 运行流水线:先检索上下文,再构造多模态请求调用LLM
retrieval_result = pipe.run({"retriever": {"query": query}})
llm_input = build_multimodal_input(
    retrieved_docs=retrieval_result["retriever"]["documents"],
    query=query,
    image_base64=image_base64
)
final_result = pipe.run({"llm": llm_input})

print(final_result["llm"]["replies"][0])

关键说明

  • 模型要求:必须使用OpenAI的多模态模型,纯文本模型(如gpt-3.5-turbo)无法处理图片输入
  • 图片格式:除了base64编码,也可以使用公开可访问的图片URL,但需确保OpenAI服务器能正常访问该地址
  • 流水线优化:如果需要更简洁的流水线结构,可以自定义一个组件封装消息构建逻辑,直接接入流水线流程

内容的提问来源于stack exchange,提问作者Abstract

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 01:50:18