使用Azure Document Intelligence分析泰卢固语PDF时输出“telugut”的修复咨询
解决Azure Document Intelligence分析泰卢固语PDF输出异常问题
针对你遇到的输出为“telugut”而非泰卢固语文本的问题,可以通过以下步骤解决:
明确指定语言参数:Azure Document Intelligence支持调用时指定目标语言,泰卢固语的语言代码为
te。你可能之前未留意该参数的设置位置:- 若使用REST API,在分析请求的JSON体中添加
"language": "te"字段; - 若使用Python SDK,在调用分析方法时传入
language="te"参数,示例代码:from azure.ai.formrecognizer import DocumentAnalysisClient from azure.core.credentials import AzureKeyCredential endpoint = "你的服务端点" key = "你的API密钥" document_analysis_client = DocumentAnalysisClient( endpoint=endpoint, credential=AzureKeyCredential(key) ) with open("泰卢固语文档.pdf", "rb") as f: poller = document_analysis_client.begin_analyze_document( "prebuilt-read", document=f, language="te" ) result = poller.result()
- 若使用REST API,在分析请求的JSON体中添加
验证API版本兼容性:确保使用Document Intelligence v3.1及以上版本,旧版本对印度区域语言的支持有限,升级到最新版本可提升识别准确性。
确认文档类型:如果你的PDF是扫描生成的图片型文档,需确保启用OCR识别模式,并同步指定
language="te"参数,避免因默认语言不匹配导致识别错误。
内容的提问来源于stack exchange,提问作者curiouslearner
相关产品推荐
相关产品推荐

