You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Document AI模型训练失败:已设可选字段仍报INVALID_DOCUMENT错误

Google Document AI Training Error: INVALID_DOCUMENT Despite Optional Field Schema

Looks like you're hitting a common gotcha with Document AI's training data validation—let's break this down clearly to fix the issue.

First, let's restate your problem for clarity: You're training a model with Google Document AI, and every document fails validation with the error below. You've tried setting the INCOME_ADJUSTMENTS field to multiple schema options (Optional Once, Optional Multiple, etc.), but the error persists.

Error Snippet

"trainingDatasetValidation": {
      "documentErrors": [
        {
          "code": 3,
          "message": "Invalid document.",
          "details": [
            {
              "@type": "type.googleapis.com/google.rpc.ErrorInfo",
              "reason": "INVALID_DOCUMENT",
              "domain": "documentai.googleapis.com",
              "metadata": {
                "num_fields": "0",
                "num_fields_needed": "1",
                "document": "5e88c5e4cc05ddb8.json",
                "annotation_name": "INCOME_ADJUSTMENTS",
                "field_name": "entities.text_anchor.text_segments"
              }
            }
          ]
        }

The Key Misunderstanding

You initially thought this error was about needing at least one INCOME_ADJUSTMENTS entity—but that's not the case. Look closely at the field_name in the metadata: entities.text_anchor.text_segments.

This error is telling you: When you include an INCOME_ADJUSTMENTS entity in your annotation, it must have at least one text segment defined in its text_anchor. The optional setting for the field only controls whether the entity needs to exist at all—not whether the entity, if present, must have valid text anchoring.

What's Likely Wrong with Your Data

  • Empty INCOME_ADJUSTMENTS entities: Some (or all) of your documents include the INCOME_ADJUSTMENTS entity in their annotation JSON, but the text_anchor object is missing text_segments entirely, or the array is empty.
  • Unintended entity entries: Even if you don't intend to annotate INCOME_ADJUSTMENTS in a document, leaving an empty entity entry for it in the JSON will trigger validation—Document AI still checks that any declared entity meets basic requirements like having a text anchor.

Fixes to Implement

1. Validate Entity Annotations

For every document that includes an INCOME_ADJUSTMENTS entity, make sure it has a valid text_anchor with at least one segment. Here's an example of a correctly formatted entity:

"entities": [
  {
    "type": "INCOME_ADJUSTMENTS",
    "textAnchor": {
      "textSegments": [
        {
          "startIndex": "45",
          "endIndex": "89"
        }
      ]
    },
    // Additional entity properties (confidence, etc.) go here
  }
]
  • The startIndex and endIndex must correspond to valid positions in the document's raw text content.

2. Remove Unintended Entity Entries

If you don't want to annotate INCOME_ADJUSTMENTS in a document, don't include the entity in the annotation JSON at all. Even an empty entity object (like {"type": "INCOME_ADJUSTMENTS"}) will trigger this validation error.

3. Double-Check Schema Alignment

Confirm that the INCOME_ADJUSTMENTS field in your schema exactly matches the entity type name in your annotations (case-sensitive, no typos). A mismatch here can cause unexpected validation behavior.

Final Note

Document AI's validation rules enforce that any annotated entity must have a valid text anchor linking it to the document content—this is non-negotiable, regardless of whether the entity is marked optional in the schema. The optional setting only controls whether the entity is required to exist in the first place.

内容的提问来源于stack exchange,提问作者Aventinus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 17:45:38