You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何无需本地存储PDF首页,直接调用AWS Textract分析?

优化PDF首页文本提取方案(避免本地文件冗余)

一、内存直接处理的优化实现

你当前方案的冗余点在于本地文件的读写,完全可以用BytesIO在内存中完成单页PDF的字节流生成,直接传给Textract,无需落地到磁盘。修改后的代码如下:

from PyPDF2 import PdfReader, PdfWriter
from io import BytesIO
import boto3

def analyse_first_page(bucket_name, file_name):
    s3 = boto3.resource("s3")
    obj = s3.Object(bucket_name, file_name)
    fs = obj.get()['Body'].read()
    pdf = PdfReader(BytesIO(fs), strict=False)
    writer = PdfWriter()
    writer.add_page(pdf.pages[0])
    
    # 直接在内存中生成字节流,替代本地文件读写
    output_stream = BytesIO()
    writer.write(output_stream)
    output_stream.seek(0)  # 将指针移到流的开头
    encoded_string = bytearray(output_stream.read())

    # 调用Textract分析文本
    textract = boto3.client('textract')
    response = textract.detect_document_text(Document={"Bytes": encoded_string})

    return response

# 调用示例
# analyse_first_page("your-bucket-name", "your-file.pdf")

二、AWS原生方案(无需手动处理PDF分页)

AWS Textract提供了异步文本检测API,支持直接从S3读取PDF文件,并且可以指定仅处理第1页,完全省去用PyPDF2拆分PDF的步骤。

实现步骤:

  1. 调用start_document_text_detection,指定S3存储位置和要处理的页码(第1页)
  2. 轮询获取检测结果(生产环境建议用SNS/SQS接收完成通知,适合批量处理场景)

代码示例:

import boto3
import time

def analyse_first_page_aws_native(bucket_name, file_name):
    textract = boto3.client('textract')
    
    # 启动异步文本检测,指定仅处理第1页
    response = textract.start_document_text_detection(
        DocumentLocation={
            'S3Object': {
                'Bucket': bucket_name,
                'Name': file_name
            }
        },
        Pages=['1']  # 指定处理第1页
    )
    
    job_id = response['JobId']
    
    # 轮询获取结果(生产环境建议用SNS/SQS触发)
    while True:
        result = textract.get_document_text_detection(JobId=job_id)
        status = result['JobStatus']
        
        if status in ['SUCCEEDED', 'FAILED']:
            break
        time.sleep(5)  # 每5秒轮询一次
    
    if status == 'SUCCEEDED':
        return result
    else:
        raise Exception(f"Textract job failed: {result.get('StatusMessage', 'Unknown error')}")

# 调用示例
# analyse_first_page_aws_native("your-bucket-name", "your-file.pdf")

方案对比:

  • 内存处理方案:适合小文件,同步返回结果,无需额外配置
  • AWS原生异步方案:适合大文件或批量处理,无需依赖第三方PDF库,由AWS负责分页和处理

内容的提问来源于stack exchange,提问作者Chukwudi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 13:01:55