如何用Python+AWS Textract实现多页OCR?求发票收据提取方案
AWS Textract调用问题修复与发票/收据提取方案
一、现有代码问题排查与修复
- 语法错误:
FeatureTypes=['TABLES'|'FORMS']中的竖线是错误写法,Python列表元素需用逗号分隔,正确格式为FeatureTypes=['TABLES', 'FORMS']。 - 异步任务未等待:
start_document_analysis是异步任务,调用后立即查询会大概率返回IN_PROGRESS状态,此时无有效分析结果,必须轮询等待任务完成。
修正后的代码示例:
import time import boto3 import logging logger = logging.getLogger(__name__) textract = boto3.client('textract') def start_analysis_job(bucket_name, document_file_name): try: response = textract.start_document_analysis( DocumentLocation={ 'S3Object': {'Bucket': bucket_name, 'Name': document_file_name}}, FeatureTypes=['TABLES', 'FORMS'], ) job_id = response['JobId'] logger.info( "Started text analysis job %s on %s.", job_id, document_file_name) except textract.exceptions.ClientError: logger.exception("Couldn't analyze text in %s.", document_file_name) raise else: return job_id def get_analysis_job(job_id): while True: try: response = textract.get_document_analysis(JobId=job_id) job_status = response['JobStatus'] logger.info("Job %s status is %s.", job_id, job_status) if job_status == 'SUCCEEDED': return response elif job_status == 'FAILED': raise Exception(f"Analysis job {job_id} failed") # 间隔5秒后再次轮询 time.sleep(5) except textract.exceptions.ClientError: logger.exception("Couldn't get data for job %s.", job_id) raise
二、发票/收据类文档专属提取方案
针对发票、收据这类结构化程度高的文档,推荐使用AWS Textract专门的AnalyzeExpense API,它内置发票收据字段识别逻辑,可直接提取商家名称、交易日期、总金额、税金额等核心字段,无需手动解析Forms/Tables结果。
1. 同步调用示例(适用于小文件)
def analyze_expense(bucket_name, document_file_name): try: response = textract.analyze_expense( Document={ 'S3Object': {'Bucket': bucket_name, 'Name': document_file_name} } ) return response except textract.exceptions.ClientError: logger.exception("Couldn't analyze expense document %s.", document_file_name) raise
2. 结果解析示例
从返回结果中提取核心字段:
def parse_expense_result(expense_response): expense_documents = expense_response['ExpenseDocuments'] parsed_data = [] for doc in expense_documents: fields = {} # 提取汇总字段 for summary_field in doc['SummaryFields']: field_type = summary_field['Type']['Text'] field_value = summary_field.get('ValueDetection', {}).get('Text', '') fields[field_type] = field_value # 提取行项目(若有) line_items = [] for line_item_group in doc.get('LineItemGroups', []): for item in line_item_group['LineItems']: item_details = {} for field in item['LineItemExpenseFields']: item_details[field['Type']['Text']] = field.get('ValueDetection', {}).get('Text', '') line_items.append(item_details) parsed_data.append({ 'summary': fields, 'line_items': line_items }) return parsed_data
3. 关键优势
- 无需手动配置FeatureTypes,API自动识别发票收据结构
- 直接返回标准化字段(如
VENDOR_NAME、TOTAL、INVOICE_DATE),减少解析工作量 - 支持多页文档,自动合并结果
内容的提问来源于stack exchange,提问作者ricardo1008
相关产品推荐
相关产品推荐

