Azure Form Recognizer在Databricks Python环境中无法识别PDF非表格文本的问题咨询
我在Databricks环境中使用Azure Cognitive Form Recognizer处理PDF文件,代码如下:
from azure.ai.formrecognizer import FormRecognizerClient from azure.core.credentials import AzureKeyCredential credential = AzureKeyCredential("aaa6123af5b843a38044538d95584c3d") endpoint= "https://myformrecognizr.cognitiveservices.azure.com/" form_recognizer_client = FormRecognizerClient(endpoint, credential) with open("/dbfs/mnt/lake/RAW/export/Picturehouse.pdf", "rb") as fd: form = fd.read() poller = form_recognizer_client.begin_recognize_content(form) form_pages = poller.result() for content in form_pages: for table in content.tables: print("Table found on page {}:".format(table.page_number)) print("Table location {}:".format(table.bounding_box)) for cell in table.cells: print("Cell text: {}".format(cell.text)) print("Location: {}".format(cell.bounding_box)) print("Confidence score: {}\n".format(cell.confidence)) if content.selection_marks: print("Selection marks found on page {}:".format(content.page_number)) for selection_mark in content.selection_marks: print("Selection mark is '{}' within bounding box '{}' and has a confidence of {}".format( selection_mark.state, selection_mark.bounding_box, selection_mark.confidence ))
目前代码能识别PDF表格内的文本(比如Item、Qty、Seat Allocation等),但无法识别PDF中的说明文本:
您可向引座员出示电子票直接进入放映厅。或者,您可在影片或活动公布的开始时间前至少15分钟到售票处取票。您需要提供预订参考号和/或支付卡来协助我们找到您的预订记录。您可点击上方的‘打印此页’链接打印本页面。
请问这是Azure Form Recognizer的设计特性,还是我的代码存在配置遗漏?
这不是代码的配置遗漏,而是你使用的API方法的设计定位导致的。
你现在用的begin_recognize_content是Form Recognizer的布局分析API,它的核心作用是提取文档里的结构化布局元素——比如表格、选择标记,还有文本行、单词的位置信息,但它不会自动把零散的文本行整合成段落,也不会专门输出非表格区域的完整说明文本(除非你自己遍历所有的content.lines或者content.words)。
如果要提取PDF里的所有文本(包括你说的那段说明),有两种更合适的方案:
1. 用Read API(最推荐)
Read API就是专门用来提取文档中所有文本内容的,不管文本在不在表格里,它都会返回所有文本行、单词,还有对应的位置和置信度。你只需要把代码里的begin_recognize_content换成begin_read_document,再调整下结果处理的逻辑就行:
# 替换后的核心代码块 poller = form_recognizer_client.begin_read_document(form) result = poller.result() for page in result: print(f"Page {page.page_number}") for line in page.lines: print(f"Line text: {line.text}") print(f"Line bounding box: {line.bounding_box}\n")
2. 用通用文档模型(General Document)
要是你既需要保留表格的结构化提取,又要拿到段落文本,可以用begin_recognize_forms调用通用文档模型,它能同时返回结构化的表格数据和非结构化的段落内容:
# 替换后的核心代码块 poller = form_recognizer_client.begin_recognize_forms(model_id="prebuilt-document") result = poller.result() for form in result: print(f"Page {form.page_number}") # 处理段落文本 for paragraph in form.paragraphs: print(f"Paragraph text: {paragraph.content}") print(f"Paragraph role: {paragraph.role}\n") # 处理表格(和你之前的逻辑兼容) for table in form.tables: print("Table found on page {}:".format(table.page_number)) for cell in table.cells: print("Cell text: {}".format(cell.content)) print("Confidence score: {}\n".format(cell.confidence))
最后给你划个重点:
begin_recognize_content:主打布局结构提取,适合拿表格、选择标记这类元素,不会主动聚合段落文本begin_read_document:主打全文本提取,适合获取文档里所有的文本内容begin_recognize_forms(通用文档模型):兼顾结构化元素和非结构化文本,适合需要同时处理表格和段落的场景
内容的提问来源于stack exchange,提问作者Patterson

