使用AWS Textract从PDF链接提取文本遇格式不支持错误求助
问题分析与解决
问题背景
需要从指定PDF链接提取文本,要求不保存到本地或S3,直接通过链接处理,但使用AWS Textract时触发UnsupportedDocumentException错误。
目标PDF链接:
pdf_url = "https://www.buelach.ch/fileadmin/files/documents/Finanzen/Bericht_zum_Budget_2023.pdf"
使用的代码:
import boto3 import requests # Specify the URL of the PDF file pdf_url = "https://www.buelach.ch/fileadmin/files/documents/Finanzen/Bericht_zum_Budget_2023.pdf" print("Get URL") # Send a request to the PDF file and get the response as a bytearray response = requests.get(pdf_url) pdf_data = bytearray(response.content) print("Get response", response) # Create a boto3 session and Textract client session = boto3.Session() textract = session.client("textract") print("Created textract") # Call Textract to detect the text in the PDF response = textract.detect_document_text(Document={"Bytes": pdf_data}) print("Get textract response") # Extract the text from the response text = "" for block in response["Blocks"]: if block["BlockType"] == "LINE": text += block["Text"] + "\n" print(text)
触发的错误:
botocore.errorfactory.UnsupportedDocumentException: An error occurred (UnsupportedDocumentException) when calling the DetectDocumentText operation: Request has unsupported document format
已尝试的字节处理方式:
pdf_data = bytearray(response.content)pdf_data = response.content
核心原因
AWS Textract的detect_document_text接口仅支持单页图像格式(如JPG、PNG),不直接支持PDF文件。要处理PDF,需选择以下两种方案之一。
解决方案
方案1:内存中将PDF转图像后识别
使用pdf2image库把PDF每一页转为图像字节流,再逐页调用detect_document_text,全程无需保存文件到本地或S3。
步骤:
- 安装依赖:
pip install pdf2image pillow
注意:
pdf2image依赖Poppler工具,Ubuntu可通过apt-get install poppler-utils安装,Windows需下载预编译包配置环境变量。
- 修改后的代码:
import boto3 import requests import io from pdf2image import convert_from_bytes pdf_url = "https://www.buelach.ch/fileadmin/files/documents/Finanzen/Bericht_zum_Budget_2023.pdf" response = requests.get(pdf_url) # 将PDF字节转为图像列表(内存处理) pages = convert_from_bytes(response.content) session = boto3.Session() textract = session.client("textract") full_text = "" for page in pages: # 将图像转为PNG字节流 img_byte_arr = io.BytesIO() page.save(img_byte_arr, format='PNG') img_bytes = img_byte_arr.getvalue() # 调用Textract识别单页图像 textract_response = textract.detect_document_text(Document={"Bytes": img_bytes}) # 提取文本 for block in textract_response["Blocks"]: if block["BlockType"] == "LINE": full_text += block["Text"] + "\n" print(full_text)
方案2:临时上传S3调用异步API
如果不想处理图像转换,可临时将PDF上传到S3,调用Textract异步API完成识别后删除S3文件,满足“不长期保存”需求。
代码示例:
import boto3 import requests import time pdf_url = "https://www.buelach.ch/fileadmin/files/documents/Finanzen/Bericht_zum_Budget_2023.pdf" response = requests.get(pdf_url) # 初始化客户端 s3 = boto3.client('s3') textract = boto3.client('textract') # 临时存储参数(替换为你的S3桶名) bucket_name = "your-temp-bucket" object_key = "temp_budget_2023.pdf" # 上传PDF到S3 s3.put_object(Bucket=bucket_name, Key=object_key, Body=response.content) # 启动异步文本检测 start_response = textract.start_document_text_detection( DocumentLocation={ 'S3Object': { 'Bucket': bucket_name, 'Name': object_key } } ) job_id = start_response['JobId'] # 轮询等待任务完成 while True: result_response = textract.get_document_text_detection(JobId=job_id) status = result_response['JobStatus'] if status in ['SUCCEEDED', 'FAILED']: break time.sleep(5) # 提取文本 full_text = "" if status == 'SUCCEEDED': for block in result_response['Blocks']: if block['BlockType'] == 'LINE': full_text += block['Text'] + "\n" # 处理分页结果(若文档页数较多) while 'NextToken' in result_response: result_response = textract.get_document_text_detection(JobId=job_id, NextToken=result_response['NextToken']) for block in result_response['Blocks']: if block['BlockType'] == 'LINE': full_text += block['Text'] + "\n" # 删除临时文件 s3.delete_object(Bucket=bucket_name, Key=object_key) print(full_text)
内容的提问来源于stack exchange,提问作者taga
相关产品推荐
相关产品推荐

