You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用AWS Textract从PDF链接提取文本遇格式不支持错误求助

问题分析与解决

问题背景

需要从指定PDF链接提取文本,要求不保存到本地或S3,直接通过链接处理,但使用AWS Textract时触发UnsupportedDocumentException错误。

目标PDF链接:

pdf_url = "https://www.buelach.ch/fileadmin/files/documents/Finanzen/Bericht_zum_Budget_2023.pdf"

使用的代码:

import boto3
import requests

# Specify the URL of the PDF file
pdf_url = "https://www.buelach.ch/fileadmin/files/documents/Finanzen/Bericht_zum_Budget_2023.pdf"
print("Get URL")

# Send a request to the PDF file and get the response as a bytearray
response = requests.get(pdf_url)
pdf_data = bytearray(response.content)
print("Get response", response)


# Create a boto3 session and Textract client
session = boto3.Session()
textract = session.client("textract")
print("Created textract")


# Call Textract to detect the text in the PDF
response = textract.detect_document_text(Document={"Bytes": pdf_data})
print("Get textract response")

# Extract the text from the response
text = ""
for block in response["Blocks"]:
    if block["BlockType"] == "LINE":
        text += block["Text"] + "\n"

print(text)

触发的错误:

botocore.errorfactory.UnsupportedDocumentException: An error occurred (UnsupportedDocumentException) when calling the DetectDocumentText operation: Request has unsupported document format

已尝试的字节处理方式:

  • pdf_data = bytearray(response.content)
  • pdf_data = response.content

核心原因

AWS Textract的detect_document_text接口仅支持单页图像格式(如JPG、PNG),不直接支持PDF文件。要处理PDF,需选择以下两种方案之一。


解决方案

方案1:内存中将PDF转图像后识别

使用pdf2image库把PDF每一页转为图像字节流,再逐页调用detect_document_text,全程无需保存文件到本地或S3。

步骤:

  1. 安装依赖:
pip install pdf2image pillow

注意:pdf2image依赖Poppler工具,Ubuntu可通过apt-get install poppler-utils安装,Windows需下载预编译包配置环境变量。

  1. 修改后的代码:
import boto3
import requests
import io
from pdf2image import convert_from_bytes

pdf_url = "https://www.buelach.ch/fileadmin/files/documents/Finanzen/Bericht_zum_Budget_2023.pdf"
response = requests.get(pdf_url)

# 将PDF字节转为图像列表(内存处理)
pages = convert_from_bytes(response.content)

session = boto3.Session()
textract = session.client("textract")

full_text = ""
for page in pages:
    # 将图像转为PNG字节流
    img_byte_arr = io.BytesIO()
    page.save(img_byte_arr, format='PNG')
    img_bytes = img_byte_arr.getvalue()
    
    # 调用Textract识别单页图像
    textract_response = textract.detect_document_text(Document={"Bytes": img_bytes})
    
    # 提取文本
    for block in textract_response["Blocks"]:
        if block["BlockType"] == "LINE":
            full_text += block["Text"] + "\n"

print(full_text)

方案2:临时上传S3调用异步API

如果不想处理图像转换,可临时将PDF上传到S3,调用Textract异步API完成识别后删除S3文件,满足“不长期保存”需求。

代码示例:

import boto3
import requests
import time

pdf_url = "https://www.buelach.ch/fileadmin/files/documents/Finanzen/Bericht_zum_Budget_2023.pdf"
response = requests.get(pdf_url)

# 初始化客户端
s3 = boto3.client('s3')
textract = boto3.client('textract')

# 临时存储参数(替换为你的S3桶名)
bucket_name = "your-temp-bucket"
object_key = "temp_budget_2023.pdf"

# 上传PDF到S3
s3.put_object(Bucket=bucket_name, Key=object_key, Body=response.content)

# 启动异步文本检测
start_response = textract.start_document_text_detection(
    DocumentLocation={
        'S3Object': {
            'Bucket': bucket_name,
            'Name': object_key
        }
    }
)
job_id = start_response['JobId']

# 轮询等待任务完成
while True:
    result_response = textract.get_document_text_detection(JobId=job_id)
    status = result_response['JobStatus']
    if status in ['SUCCEEDED', 'FAILED']:
        break
    time.sleep(5)

# 提取文本
full_text = ""
if status == 'SUCCEEDED':
    for block in result_response['Blocks']:
        if block['BlockType'] == 'LINE':
            full_text += block['Text'] + "\n"
    # 处理分页结果(若文档页数较多)
    while 'NextToken' in result_response:
        result_response = textract.get_document_text_detection(JobId=job_id, NextToken=result_response['NextToken'])
        for block in result_response['Blocks']:
            if block['BlockType'] == 'LINE':
                full_text += block['Text'] + "\n"

# 删除临时文件
s3.delete_object(Bucket=bucket_name, Key=object_key)

print(full_text)

内容的提问来源于stack exchange,提问作者taga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 19:14:58