You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+AWS Textract实现多页OCR?求发票收据提取方案

AWS Textract调用问题修复与发票/收据提取方案

一、现有代码问题排查与修复

  1. 语法错误:FeatureTypes=['TABLES'|'FORMS'] 中的竖线是错误写法,Python列表元素需用逗号分隔,正确格式为 FeatureTypes=['TABLES', 'FORMS']。
  2. 异步任务未等待:start_document_analysis 是异步任务,调用后立即查询会大概率返回IN_PROGRESS状态,此时无有效分析结果,必须轮询等待任务完成。

修正后的代码示例:

import time
import boto3
import logging

logger = logging.getLogger(__name__)
textract = boto3.client('textract')

def start_analysis_job(bucket_name, document_file_name):
    try:
        response = textract.start_document_analysis(
            DocumentLocation={
                'S3Object': {'Bucket': bucket_name, 'Name': document_file_name}},
            FeatureTypes=['TABLES', 'FORMS'],
        )
        job_id = response['JobId']
        logger.info(
            "Started text analysis job %s on %s.", job_id, document_file_name)
    except textract.exceptions.ClientError:
        logger.exception("Couldn't analyze text in %s.", document_file_name)
        raise
    else:
        return job_id

def get_analysis_job(job_id):
    while True:
        try:
            response = textract.get_document_analysis(JobId=job_id)
            job_status = response['JobStatus']
            logger.info("Job %s status is %s.", job_id, job_status)
            
            if job_status == 'SUCCEEDED':
                return response
            elif job_status == 'FAILED':
                raise Exception(f"Analysis job {job_id} failed")
            
            # 间隔5秒后再次轮询
            time.sleep(5)
        except textract.exceptions.ClientError:
            logger.exception("Couldn't get data for job %s.", job_id)
            raise

二、发票/收据类文档专属提取方案

针对发票、收据这类结构化程度高的文档,推荐使用AWS Textract专门的AnalyzeExpense API,它内置发票收据字段识别逻辑,可直接提取商家名称、交易日期、总金额、税金额等核心字段,无需手动解析Forms/Tables结果。

1. 同步调用示例(适用于小文件)

def analyze_expense(bucket_name, document_file_name):
    try:
        response = textract.analyze_expense(
            Document={
                'S3Object': {'Bucket': bucket_name, 'Name': document_file_name}
            }
        )
        return response
    except textract.exceptions.ClientError:
        logger.exception("Couldn't analyze expense document %s.", document_file_name)
        raise

2. 结果解析示例

从返回结果中提取核心字段:

def parse_expense_result(expense_response):
    expense_documents = expense_response['ExpenseDocuments']
    parsed_data = []
    
    for doc in expense_documents:
        fields = {}
        # 提取汇总字段
        for summary_field in doc['SummaryFields']:
            field_type = summary_field['Type']['Text']
            field_value = summary_field.get('ValueDetection', {}).get('Text', '')
            fields[field_type] = field_value
        
        # 提取行项目(若有)
        line_items = []
        for line_item_group in doc.get('LineItemGroups', []):
            for item in line_item_group['LineItems']:
                item_details = {}
                for field in item['LineItemExpenseFields']:
                    item_details[field['Type']['Text']] = field.get('ValueDetection', {}).get('Text', '')
                line_items.append(item_details)
        
        parsed_data.append({
            'summary': fields,
            'line_items': line_items
        })
    
    return parsed_data

3. 关键优势

  • 无需手动配置FeatureTypes,API自动识别发票收据结构
  • 直接返回标准化字段(如VENDOR_NAME、TOTAL、INVOICE_DATE),减少解析工作量
  • 支持多页文档,自动合并结果

内容的提问来源于stack exchange,提问作者ricardo1008

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 12:01:02