在Lambda函数中批量传递S3文件至AWS Textract API提速咨询
问题与优化方案
问题背景
我有一张包含多文本字段的PDF图片,已裁剪出对应每个文本框的小图,用字典管理各图片对应的字段。原代码通过串行调用AWS Textract API提取文本,处理数百张图耗时约7分钟,需要缩短至45秒内。尝试过传入S3文件名列表但报错,不清楚如何实现并行处理或批量提交并保持字段对应关系。
原字典结构示例:
aws_output = { "Section A": {"1": extract_text(imgs['Section A']['1'], "sa1"), "2": extract_text(imgs['Section A']['2'], "sa2")}, "Section B": {"1": extract_text(imgs['Section B']['1'], "sb1")} }
原提取函数:
def extract_text(img, loc): img_obj = Image.fromarray(img).convert('RGB') out_img_obj = io.BytesIO() img_obj.save(out_img_obj, format="png") out_img_obj.seek(0) file_name = key_id + "_" + loc + ".png" s3.Bucket(bucket_name).put_object(Key=file_name, Body=out_img_obj, ContentType="image/png") response = textract_client.detect_document_text( Document={ 'S3Object': { 'Bucket': bucket_name, 'Name': file_name } } ) status_code = response['ResponseMetadata']['HTTPStatusCode'] try: status_code == 200 text_len = {} for y in range(len(response['Blocks'])): if 'Text' in response['Blocks'][y]: text_len[y] = len(response['Blocks'][y]['Text']) else: pass if bool(text_len): extracted_text = response['Blocks'][max(text_len, key=text_len.get)]['Text'] if extracted_text == '-': extracted_text = '' else: pass else: extracted_text = '' s3.Object(bucket_name,file_name).delete() return extracted_text except: return f"HTTP Status code is not 200 when running extract_text on {loc}, code is {status_code}"
优化方案
1. 跳过S3中转,直接传递图片字节
Textract的detect_document_text支持直接传入图片字节(Bytes参数),无需先上传到S3再删除,这能节省大量IO时间。修改提取函数去掉S3相关操作:
def extract_text(img, loc): try: # 转换图片为字节流 img_obj = Image.fromarray(img).convert('RGB') out_img_obj = io.BytesIO() img_obj.save(out_img_obj, format="png") out_img_obj.seek(0) img_bytes = out_img_obj.getvalue() # 直接调用Textract,传入字节流 response = textract_client.detect_document_text( Document={'Bytes': img_bytes} ) # 提取最长文本(保留原逻辑) text_len = {} for block in response['Blocks']: if 'Text' in block: text_len[block['Id']] = len(block['Text']) if text_len: extracted_text = response['Blocks'][max(text_len, key=text_len.get)]['Text'] return '' if extracted_text == '-' else extracted_text return '' except Exception as e: return f"Error extracting text for {loc}: {str(e)}"
2. 用多线程并行处理
由于调用Textract和图片转换属于IO密集型操作,使用多线程可以大幅提升效率。用concurrent.futures.ThreadPoolExecutor实现并行,同时打包每个任务的上下文(section、key、img、loc),确保结果能正确映射回原字典:
from concurrent.futures import ThreadPoolExecutor def process_all_images(imgs): # 整理所有任务:每个任务包含section、key、img、loc tasks = [] for section, items in imgs.items(): for key, img in items.items(): loc = f"{section.lower().replace(' ', '')}{key}" # 生成类似sa1的loc tasks.append( (section, key, img, loc) ) # 并行执行,设置线程数(可根据AWS配额调整,建议20-50) aws_output = {section: {} for section in imgs.keys()} with ThreadPoolExecutor(max_workers=30) as executor: # 提交所有任务,保留future与上下文的映射 future_to_task = {executor.submit(extract_text, img, loc): (section, key) for section, key, img, loc in tasks} # 逐个获取结果并填充字典 for future in future_to_task: section, key = future_to_task[future] try: result = future.result() aws_output[section][key] = result except Exception as e: aws_output[section][key] = f"Task failed: {str(e)}" return aws_output # 使用示例 aws_output = process_all_images(imgs)
3. 注意事项
- AWS Textract有调用配额,默认同步调用是1000次/秒,线程数不要超过配额限制,避免触发限流。如果需要更高并发,可以申请提升配额。
- 若图片数量极大(数千张),可以考虑使用Textract的异步批量处理(
StartDocumentTextDetection),但需要处理任务队列、轮询结果,复杂度稍高,适合超大规模场景。 - 原代码中
status_code == 200是无效判断,改为直接捕获异常即可,因为调用失败时boto3会直接抛出异常。
内容的提问来源于stack exchange,提问作者carousallie
相关产品推荐
相关产品推荐

