关于Google Document AI无法返回textStyle属性的技术咨询
Google Document AI textStyle属性为空问题说明
textStyle属性是Document AI正式支持的功能,并非仅文档展示用途,返回为空通常由以下原因导致:
1. 处理器类型不匹配
只有特定类型的处理器会返回文本样式信息,基础OCR处理器不支持该特性。需在Cloud Console中创建并使用支持文本样式提取的处理器,例如:
- 进阶版OCR处理器(带样式提取的OCR Processor v1)
- 表单解析器(Form Parser)
- 文档理解类处理器
2. 未正确遍历文档结构获取textStyle
你提供的示例代码仅打印了文档纯文本内容,而textStyle属性存储在文档的层级结构中(页面→块→段落→单词),需要逐层遍历才能获取。修改代码如下:
from google.api_core.client_options import ClientOptions from google.cloud import documentai_v1 as documentai def quickstart( project_id: str, location: str, processor_id: str, file_path: str, mime_type: str ): opts = ClientOptions(api_endpoint=f"{location}-documentai.googleapis.com") client = documentai.DocumentProcessorServiceClient(client_options=opts) name = client.processor_path(project_id, location, processor_id) with open(file_path, "rb") as image: image_content = image.read() raw_document = documentai.RawDocument(content=image_content, mime_type=mime_type) request = documentai.ProcessRequest(name=name, raw_document=raw_document) result = client.process_document(request=request) document = result.document # 遍历文档结构提取textStyle for page in document.pages: for block in page.blocks: for paragraph in block.paragraphs: for word in paragraph.words: text_segment = document.text[word.start_index:word.end_index] text_style = word.text_style print(f"文本内容: {text_segment}") print(f"字体大小: {text_style.font_size}") print(f"加粗: {text_style.bold}") print(f"斜体: {text_style.italic}") print(f"下划线: {text_style.underline}") print("---")
补充说明
- 扫描件(图像PDF)的样式信息由OCR识别生成,可能存在一定误差,但会正常返回;机器生成PDF的样式提取精度更高。
- 确保使用最新版本的Google Cloud Document AI Python客户端,避免因版本兼容问题导致属性缺失。
内容的提问来源于stack exchange,提问作者Ankit A
相关产品推荐
相关产品推荐

