You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Amazon Bedrock Converse API提交图片遇ValidationException问题求助

解决Amazon Bedrock Converse接口图片处理ValidationException问题

问题场景

使用boto3调用Amazon Bedrock Runtime的Converse接口提交文本+Base64图片时,遇到以下错误:

ValidationException: An error occurred (ValidationException) when calling the Converse operation: The model returned the following errors: Unable to process provided image.

代码片段如下:

import boto3
import json

config = {
    "additionalModelRequestFields": {},
    "inferenceConfig": {
        "maxTokens": 512,
        "temperature": 0.5,
        "topP": 0.9
    },
    "messages": [
        {
            "content": [
                {"text": "What is this document?"},
                {"image": {"format": "jpeg", "source": {"bytes": img_base64_str}}}
            ],
            "role": "user"
        }
    ]
}

client = boto3.client('bedrock-runtime', region_name='us-east-1', aws_access_key_id=aws_access_key_id, aws_secret_access_key=aws_secret_access_key)

model_id = 'us.meta.llama3-2-11b-instruct-v1:0'

response = client.converse(inferenceConfig=config['inferenceConfig'], messages=config['messages'], modelId=model_id)

已确认Base64字符串有效且更换过不同图片,错误仍存在,需明确图片输入要求及排查方案。

核心原因与解决方案

1. 模型不支持多模态输入

你使用的us.meta.llama3-2-11b-instruct-v1:0是纯文本指令模型,不具备视觉处理能力。需要更换为Llama3-2的多模态版本,比如:

  • us.meta.llama3-2-11b-vision-v1:0
  • us.meta.llama3-2-7b-vision-v1:0

2. 图片输入格式要求

若已使用支持多模态的模型,需满足以下要求:

  • Base64字符串需纯净:不能带data:image/jpeg;base64,这类前缀,必须是原始图片二进制数据直接编码的字符串。
  • 格式匹配:请求中指定的format(如jpeg)必须与图片实际格式完全一致,比如PNG图片需将format设为png。
  • 尺寸与大小限制:不同模型有不同限制,以Llama3-2 Vision为例,通常要求图片分辨率不超过4096x4096,文件大小不超过5MB,超出需压缩。

3. 修正后的代码示例

import boto3
import json
import base64

# 读取图片并生成纯净Base64字符串
with open("your-image.jpg", "rb") as f:
    img_base64_str = base64.b64encode(f.read()).decode("utf-8")

config = {
    "inferenceConfig": {
        "maxTokens": 512,
        "temperature": 0.5,
        "topP": 0.9
    },
    "messages": [
        {
            "content": [
                {"text": "What is this document?"},
                {"image": {"format": "jpeg", "source": {"bytes": img_base64_str}}}
            ],
            "role": "user"
        }
    ]
}

client = boto3.client('bedrock-runtime', region_name='us-east-1')

# 使用支持多模态的模型ID
model_id = 'us.meta.llama3-2-11b-vision-v1:0'

response = client.converse(inferenceConfig=config['inferenceConfig'], messages=config['messages'], modelId=model_id)
print(response)

排查步骤

  • 先确认模型是否支持视觉输入:在Bedrock控制台查看模型详情,确认标注有"Multimodal"或"Vision"能力。
  • 验证Base64有效性:将字符串解码为二进制文件,检查是否能正常打开图片。
  • 核对图片格式:用图片查看工具确认实际格式,确保与请求中的format参数一致。
  • 压缩图片:若图片过大,使用工具压缩到模型允许的大小范围内。

内容的提问来源于stack exchange,提问作者Nathan Bennett

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 20:22:45