如何使用Google Cloud Vision读取多页PDF并解决代码报错与单页识别问题
问题背景
当前尝试使用Google Cloud Vision API读取多页PDF文件,仅能读取PDF的第一页,同时代码运行触发报错,附相关物料如下:

报错修复方案
你遇到的索引越界报错是因为代码硬编码取下标为0的响应元素,当请求返回的响应列表为空、或返回的结构不符合预期时就会触发该问题。
你需要将代码中直接取response.responses[0]的逻辑替换为遍历所有响应元素,示例修改如下:
# 原错误写法(示例) # page_text = response.responses[0].responses[0].full_text_annotation.text # 修正后的写法 full_text = "" for file_resp in response.responses: for page_resp in file_resp.responses: full_text += page_resp.full_text_annotation.text + "\f" # 用换页符分隔不同页面内容
多页PDF完整读取实现
Google Cloud Vision API默认仅处理PDF的第一页,要读取全页内容需要在构造请求时新增pages参数指定要处理的页码:
- 先统计待处理PDF的总页数,如果你不确定总页数可以用PyPDF2等工具提前读取
- 构造请求时添加
pages字段,单次请求最多支持处理100页,超过100页需要分批请求
示例请求构造代码:
from google.cloud import vision client = vision.ImageAnnotatorClient() # 配置输入输出 gcs_source = vision.GcsSource(uri="gs://你的桶地址/目标文件.pdf") input_config = vision.InputConfig(gcs_source=gcs_source, mime_type="application/pdf") gcs_destination = vision.GcsDestination(uri="gs://你的桶地址/输出路径/") output_config = vision.OutputConfig(gcs_destination=gcs_destination, batch_size=100) # 这里替换为你实际的PDF总页数,单次请求最多填100 total_pages = 50 request = vision.AsyncAnnotateFileRequest( features=[{"type_": vision.Feature.Type.DOCUMENT_TEXT_DETECTION}], input_config=input_config, output_config=output_config, pages=list(range(1, total_pages + 1)) # 关键:指定要处理的所有页码 ) # 发起异步请求 operation = client.async_batch_annotate_files(requests=[request]) response = operation.result(timeout=600) # 合并所有页面文本 full_pdf_text = "" for file_response in response.responses: for page_response in file_response.responses: full_pdf_text += page_response.full_text_annotation.text + "\n--- 分页符 ---\n"
如果PDF页数超过100,把页码列表拆分为多个长度为100的子列表,依次发起请求再合并所有结果即可。
内容的提问来源于stack exchange,提问作者12345
相关产品推荐
相关产品推荐

