SageMaker调用create_labeling_job()生成标注结果异常求助
问题:SageMaker Bounding Box标注任务输出无有效数据
通过AWS控制台创建图像Bounding Box标注任务可正常完成,但使用boto3的create_labeling_job()函数创建的任务,虽状态显示COMPLETE、标注已完成,输出manifest文件却无有效标注数据:image_size宽高为0,annotations为空。
用户提供的相关文件
1. Boto3代码示例
task_description = "Draw bounding box around the images" task_keywords = ['Images', 'bounding boxes', 'object detection'] task_title = "Bounding Box task" job_name = "test-labeling-job" acs_arn = "arn:aws:lambda:eu-west-1:568282634449:function:ACS-BoundingBox" prehuman_arn = "arn:aws:lambda:eu-west-1:568282634449:function:PRE-BoundingBox" human_task_config = { "AnnotationConsolidationConfig": { "AnnotationConsolidationLambdaArn": acs_arn, }, "PreHumanTaskLambdaArn": prehuman_arn, "MaxConcurrentTaskCount": 1000, # 200 images will be sent at a time to the workteam. "NumberOfHumanWorkersPerDataObject": 1, # 3 separate workers will be required to label each image. "TaskAvailabilityLifetimeInSeconds": 864000, # Your workteam has 6 hours to complete all pending tasks. "TaskDescription": task_description, "TaskKeywords": task_keywords, "TaskTimeLimitInSeconds": 3600, # Each image must be labeled within 5 minutes. "TaskTitle": task_title, "UiConfig": { "UiTemplateS3Uri": "s3://bucket-name/bounding-box.liquid.html", }, "WorkteamArn": "arn:aws:sagemaker:eu-west-1:xxxxxxxxx:workteam/private-crowd/Labellers" } ground_truth_request = { "InputConfig": { "DataSource": { "S3DataSource": { "ManifestS3Uri": "s3://bucket-name/manifest-vic-2.manifest", } }, "DataAttributes": { "ContentClassifiers": [] }, }, "OutputConfig": { "S3OutputPath": "s3://sagemaker-bucket-name/folder-name-07-12-2022/", }, "StoppingConditions":{ 'MaxPercentageOfInputDatasetLabeled': 100 }, "HumanTaskConfig": human_task_config, "LabelingJobName": job_name, "RoleArn": "arn:aws:iam::xxxxxxxxx:role/service-role/AmazonSageMaker-ExecutionRole-xxxxxxxxxx", "LabelAttributeName": job_name, "LabelCategoryConfigS3Uri": "s3://bucket-name/label-config.json", } sagemaker_client = boto3.client("sagemaker") sagemaker_client.create_labeling_job(**ground_truth_request)
2. 标签配置文件(label-config.json)
{ "document-version": "2018-11-28", "labels": [{"label": "animal"}, {"label": "no-animal"}] }
3. 输入manifest文件
{'source-ref': 's3://sagemaker-inout-folder/2022-10-10-12:50:59/A.jpg'} {'source-ref': 's3://sagemaker-inout-folder/2022-10-10-12:50:59/B.jpg'} {'source-ref': 's3://sagemaker-inout-folder/2022-10-10-12:50:59/C.jpg'}
4. UI模板文件(bounding-box.liquid.html)
<script src="https://assets.crowd.aws/crowd-html-elements.js"></script> <crowd-form> <crowd-bounding-box name="annotatedResult" src="{{ task.input.taskObject | grant_read_access }}" header="Draw bounding boxes around all the cats and dogs in this image" labels="['Animal', 'No-animal']" > <full-instructions header="Bounding Box Instructions" > <p>Use the bounding box tool to draw boxes around the requested target of interest:</p> <ol> <li>Draw a rectangle using your mouse over each instance of the target.</li> <li>Make sure the box does not cut into the target, leave a 2 - 3 pixel margin</li> <li> When targets are overlapping, draw a box around each object, include all contiguous parts of the target in the box. Do not include parts that are completely overlapped by another object. </li> <li> Do not include parts of the target that cannot be seen, even though you think you can interpolate the whole shape of the target. </li> <li>Avoid shadows, they're not considered as a part of the target.</li> <li>If the target goes off the screen, label up to the edge of the image.</li> </ol> </full-instructions> <short-instructions> Draw boxes around the requested target of interest. </short-instructions> </crowd-bounding-box> </crowd-form>
5. 输出manifest示例
{ "source-ref": "s3://sagemaker-inout-folder/2022-10-10-12:50:59/A.jpg", "test-labeling-job": { "image_size": [ { "width": 0, "height": 0, "depth": 3 } ], "annotations": [] }, "test-labeling-job-metadata": { "objects": [], "class-map": {}, "type": "groundtruth/object-detection", "human-annotated": "yes", "creation-date": "2022-12-01T08:09:07.933647", "job-name": "labeling-job/test-labeling-job" } }
问题排查与解决方法
1. 修正Lambda函数ARN(核心问题)
你使用的Pre/聚合Lambda是自定义ARN,若未正确适配Bounding Box任务逻辑,会导致图片尺寸无法读取、标注无法聚合。建议替换为eu-west-1区域的官方托管Lambda:
- 预任务Lambda:
arn:aws:lambda:eu-west-1:432418664414:function:PRE-BoundingBox - 聚合任务Lambda:
arn:aws:lambda:eu-west-1:432418664414:function:ACS-BoundingBox
若坚持使用自定义Lambda,需检查函数是否正确返回图片尺寸、是否能解析标注结果并生成符合要求的输出结构。
2. 统一标签大小写
标签配置文件中标签为全小写(animal/no-animal),但UI模板中标签首字母大写(Animal/No-animal),大小写不匹配会导致标注无法映射到标签体系,最终annotations为空。修改UI模板的labels字段:
labels="['animal', 'no-animal']"
3. 修正输入Manifest格式
输入Manifest必须是标准JSON格式,需将单引号替换为双引号:
{"source-ref": "s3://sagemaker-inout-folder/2022-10-10-12:50:59/A.jpg"} {"source-ref": "s3://sagemaker-inout-folder/2022-10-10-12:50:59/B.jpg"} {"source-ref": "s3://sagemaker-inout-folder/2022-10-10-12:50:59/C.jpg"}
4. 检查SageMaker执行角色权限
确保执行角色拥有以下权限:
- 输入/输出S3桶的读写权限(
s3:GetObject/s3:PutObject) - 调用指定Lambda函数的权限(
lambda:InvokeFunction) - 访问图片存储桶的权限(若图片在独立桶中)
5. 验证UI模板变量映射
确认UI模板中的{{ task.input.taskObject | grant_read_access }}正确映射到输入Manifest的source-ref字段,自定义Pre函数需确保返回结构包含taskObject字段,否则无法加载图片。
内容的提问来源于stack exchange,提问作者impromptu_user
相关产品推荐
相关产品推荐

