使用Document AI自定义提取器自动标注时的Schema要求问题
解决Document AI Custom Extractor自动标注时的"No valid schema provided"错误
问题场景
使用Google Cloud Document AI的Custom Extractor创建自定义提取器时,启用Auto-label自动标注功能(选择唯一可用版本pretrained-foundation-model-v1.0-2023-08-22)上传142份文档后,触发如下错误:
{ "name": "projects/xxxxxxxxx/locations/xxxxxxx/operations/xxxxxxx", "done": true, "result": "error", "response": {}, "metadata": { "@type": "type.googleapis.com/google.cloud.documentai.uiv1beta3.ImportDocumentsMetadata", "commonMetadata": { "state": "FAILED", "createTime": "202x-xxx-xxT01:xx:45.367220Z", "updateTime": "202x-xxx-xxT01:xx:57.243001Z", "resource": "projects/xxxxxxx/locations/xxxxxx/processors/xxxxxxxxx/dataset" }, "totalDocumentCount": 142 }, "error": { "code": 3, "message": "No valid schema provided for processing.", "details": [] } }
解决步骤
先创建并关联基础Schema
Auto-label功能无法完全从零生成标签体系,必须依赖一个基础Schema作为框架。即使你希望系统自动生成所有标签,也要先创建一个最小化的Schema:- 进入数据集管理的「Schema」选项卡
- 创建新Schema,至少添加1个合法字段(比如命名为
temp_field,类型选文本即可) - 确保Schema状态为「已启用」并关联到当前数据集
检查Schema的有效性
如果已有Schema仍报错,排查以下细节:- 字段名不能包含特殊字符(如
!@#$%),不能为空 - 所有字段必须指定明确的类型(文本、日期、数值等,不能选「未知」)
- 确认Schema的区域与数据集区域一致(比如都在
us-central1) - 确保Schema已设置为「当前活跃版本」
- 字段名不能包含特殊字符(如
重新执行上传流程
- 取消当前失败的上传任务,清理临时文档
- 确认Schema已正确关联到数据集后,重新勾选Auto-label选项
- 选择指定的基础模型版本,再次上传文档
内容的提问来源于stack exchange,提问作者tmighty
相关产品推荐
相关产品推荐

