使用JavaScript的ConverseCommand向AWS Bedrock发送图片失败求助
问题:使用AWS Bedrock Runtime ConverseCommand上传图片失败
我使用@aws-sdk/client-bedrock-runtime v3.686版本的ConverseCommand,在对话场景下向AWS Bedrock上传图片。按照官方文档要求,需传入Uint8Array格式的字节数据,我认为这和Node.js的Buffer等价,但尝试了.buffer、new TextEncoder().encode(imageBuffer)等多种转换方式后,始终收到模型提示:
I'm sorry, but you haven't actually shared an image with me yet. Could you please upload an image and I'll take a look at it?
找不到包含图片上传场景的@aws-sdk/client-bedrock-runtime可用示例代码,希望得到解决思路。
原尝试代码
const fs = require("fs"); const { BedrockRuntimeClient, ConverseCommand, } = require("@aws-sdk/client-bedrock-runtime"); const modelId = "anthropic.claude-3-haiku-20240307-v1:0"; let conversation = []; const askQuestion = async () => { const client = new BedrockRuntimeClient({ region: "us-east-1", credentials: { accessKeyId: process.env.AWS_ACCESS_KEY_ID, secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY, }, }); let content = {}; const imageBuffer = fs.readFileSync("./butterfly.png"); content.image = { format: "png", source: { bytes: imageBuffer.buffer, }, }; content.text = "What's in this image?"; conversation.push({ role: "user", content: [content], }); try { const response = await client.send( new ConverseCommand({ modelId, messages: conversation, }) ); conversation.push(response.output.message); console.log(response.output.message.content[0].text); return response.output.message.content[0].text; } catch (err) { console.log(`ERROR: Can't invoke '${modelId}'. Reason: ${err}`); } }; askQuestion();
修正后可正常运行的代码
经调整图片数据的转换方式和内容结构后,代码可正常调用模型识别图片:
const fs = require("fs"); const { BedrockRuntimeClient, ConverseCommand, } = require("@aws-sdk/client-bedrock-runtime"); const modelId = "anthropic.claude-3-haiku-20240307-v1:0"; let conversation = []; const askQuestion = async () => { const client = new BedrockRuntimeClient({ region: "us-east-1", credentials: { accessKeyId: process.env.AWS_ACCESS_KEY_ID, secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY, }, }); let content = []; const imageBuffer = new Uint8Array(fs.readFileSync("./butterfly.png").buffer); content.push({ image: { format: "png", source: { bytes: imageBuffer, }, }, }); content.push({ text: "What's in this image?" }); conversation.push({ role: "user", content, }); try { const response = await client.send( new ConverseCommand({ modelId, messages: conversation, }) ); conversation.push(response.output.message); console.log(response.output.message.content[0].text); return response.output.message.content[0].text; } catch (err) { console.log(`ERROR: Can't invoke '${modelId}'. Reason: ${err}`); } }; askQuestion();
关键修正点
- 将
content从对象改为数组,分别添加image和text类型的内容项,符合多模态消息的结构要求 - 正确将图片Buffer转换为Uint8Array:
new Uint8Array(fs.readFileSync("./butterfly.png").buffer)
内容的提问来源于stack exchange,提问作者Jeff Bonnes
相关产品推荐
相关产品推荐

