You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure.AI.DocumentIntelligence .NET SDK自定义模型多页分析异常排查

问题:Azure.AI.DocumentIntelligence SDK分析多页文档仅返回单页结果

我正在使用Azure.AI.DocumentIntelligence SDK for .NET,基于自定义模型分析多页文档,但遇到问题:返回的AnalyzeResult对象仅包含单页结果,尽管文档表明该SDK支持多页分析。

我的代码

private async Task<AnalyzeResult> ExtractDocumentAsync(string filePath, string modelId)
{
    // Your Form Recognizer endpoint and API key
    string endpoint = system.GetAsset("AzureDU_Endpoint","Root/.General").ToString();
    string apiKey = system.GetAsset("AzureDU_ApiKey","Root/.General").ToString();
    
    var client = new DocumentIntelligenceClient(new Uri(endpoint), new  AzureKeyCredential(apiKey));
    // Now you can use the client to interact with the Form Recognizer service
    // For example:
        try
        {
            //string documentText = system.ReadTextFile(system.GetResourceForLocalPath(filePath,PathType.File));
            //string fileContent = Convert.ToBase64String(Encoding.UTF8.GetBytes(documentText));
            byte[] fileContent = File.ReadAllBytes(filePath);
            //Uri uriSource = new Uri("file://" + filePath);
            var content = new AnalyzeDocumentContent()
            {
                 Base64Source= System.BinaryData.FromBytes(fileContent)
                //UrlSource=uriSource
                //Base64Source = BinaryData.FromBytes(Encoding.UTF8.GetBytes(fileContent))

            };
            
            Operation<AnalyzeResult> operation = await client.AnalyzeDocumentAsync(WaitUntil.Completed, modelId, content,"1,2,3");
            AnalyzeResult result = operation.Value;
            Console.WriteLine($"Document was analyzed with model with ID: {result.ModelId}");
            return result;
        }
        catch (Exception ex)
        {
            Console.WriteLine($"An error occurred: {ex.Message}");
            throw;
        }
}

解决方案

问题出在调用AnalyzeDocumentAsync时的页码参数格式错误:你传入的"1,2,3"不符合服务要求的页码格式,服务无法正确解析逗号分隔的页码列表,导致仅返回第一页结果。

服务支持的页码参数格式为:

  • 单个页码:如"1"
  • 连续页码范围:如"1-3"(表示分析第1到第3页)
  • 若要分析所有页面,直接省略该参数即可

修改后的关键代码如下:

// 方式1:分析文档所有页面
Operation<AnalyzeResult> operation = await client.AnalyzeDocumentAsync(WaitUntil.Completed, modelId, content);

// 方式2:指定连续页码范围(如第1到第3页)
// Operation<AnalyzeResult> operation = await client.AnalyzeDocumentAsync(WaitUntil.Completed, modelId, content, "1-3");

另外还需确认两个前提:

  • 你上传的文档确实是多页格式(如PDF、TIFF等),而非单页文档
  • 自定义模型的训练不影响多页分析的返回,即使模型基于单页数据训练,服务仍会返回所有页面的结果,只是未训练的页面字段识别精度可能较低

内容的提问来源于stack exchange,提问作者יונתן מנגיסטו

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 05:11:35