You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Lambda生成Parquet成功却超时,求解决方案

问题

我正在开发一款基于Python的AWS Lambda函数,功能为将JSON文件转换为Parquet格式。已设置5分钟超时限制,调用函数后发现Parquet文件已成功生成,但函数仍触发超时。调试发现程序在执行to_parquet代码行时陷入停滞,而该函数在本地环境可正常运行。相关代码及AWS Lambda输出日志如下:

Python函数代码

import awswrangler as wr
import pandas as pd
import urllib.parse
import os

# Temporary hard-coded AWS Settings; i.e. to be set as OS variable in Lambda
os_input_s3_cleansed_layer = os.environ['s3_cleansed_layer']
os_input_glue_catalog_db_name = os.environ['glue_catalog_db_name']
os_input_glue_catalog_table_name = os.environ[ 'glue_catalog_table_name']
os_input_write_data_operation = os.environ['write_data_operation']

def lambda_handler (event, context):
    print('## EVENT')
    print(event)
    # Get the object from the event and show its content type
    bucket = event['Records'][0]['s3']['bucket']['name']
    key = urllib.parse.unquote_plus(event['Records'][0]['s3']['object']['key'], encoding='utf-8')
    try:
        # Creating DF from content
        df_raw = wr.s3.read_json ('s3://{}/{}'.format(bucket, key))
        print('## DF RAW')
        print(df_raw.shape)
        # Extract required columns:
        df_step_1 = pd.json_normalize(df_raw['items'])
        print('## DF CLEANED')
        print(df_step_1.shape)
        # Write to S3
        wr_response = wr.s3.to_parquet(
            df=df_step_1,
            path=os_input_s3_cleansed_layer,
            dataset=True,
            database=os_input_glue_catalog_db_name,
            table=os_input_glue_catalog_table_name,
            mode=os_input_write_data_operation
        )
        print('## RESPONSE')
        return wr_response
    except Exception as e:
        print(e)
        print('Error getting object {} from bucket {}.Make sure they exist and your bucket is in the same region as this function.'.format (key, bucket))
        raise e

AWS Lambda输出日志

Test Event Name
s3-put

Response
{
  "errorMessage": "2024-07-31T07:11:36.885Z 7a24ab3f-13c8-4284-980f-6b705d64273a Task timed out after 307.11 seconds"
}

Function Logs
START RequestId: 7a24ab3f-13c8-4284-980f-6b705d64273a Version: $LATEST
## EVENT
{'Records': [{'eventVersion': '2.0', 'eventSource': 'aws:s3', 'awsRegion': 'us-east-1', 'eventTime': '1970-01-01T00:00:00.000Z', 'eventName': 'ObjectCreated:Put', 'userIdentity': {'principalId': 'EXAMPLE'}, 'requestParameters': {'sourceIPAddress': '127.0.0.1'}, 'responseElements': {'x-amz-request-id': 'EXAMPLE123456789', 'x-amz-id-2': 'EXAMPLE123/5678abcdefghijklambdaisawesome/mnopqrstuvwxyzABCDEFGH'}, 's3': {'s3SchemaVersion': '1.0', 'configurationId': 'testConfigRule', 'bucket': {'name': 'youtubepipeline-raw-useast1-dev', 'ownerIdentity': {'principalId': 'EXAMPLE'}, 'arn': 'arn:aws:s3:::youtubepipeline-raw-useast1-dev'}, 'object': {'key': 'youtube/raw_statistics_reference_data/CA_category_id.json', 'size': 1024, 'eTag': '0123456789abcdef0123456789abcdef', 'sequencer': '0A1B2C3D4E5F678901'}}}]} 
## DF RAW
(31, 3)
## DF CLEANED
(31, 6)
2024-07-31T07:11:36.885Z 7a24ab3f-13c8-4284-980f-6b705d64273a Task timed out after 307.11 seconds

END RequestId: 7a24ab3f-13c8-4284-980f-6b705d64273a
REPORT RequestId: 7a24ab3f-13c8-4284-980f-6b705d64273a  Duration: 307106.33 ms  Billed Duration: 300000 ms  Memory Size: 128 MB Max Memory Used: 128 MB Init Duration: 4043.97 ms

Request ID
7a24ab3f-13c8-4284-980f-6b705d64273a
解决建议
  • 提升Lambda内存配置:日志显示函数已耗尽128MB内存。AWS Lambda的CPU、网络带宽与内存配置正相关,将内存提升至256MB或512MB,可显著加快数据处理和API调用速度,避免资源不足导致的停滞。
  • 检查Glue Catalog权限与区域一致性:启用dataset=True后,to_parquet会自动同步Glue元数据。若Lambda执行角色缺少Glue的GetTable/UpdateTable/CreateTable权限,或Glue Catalog与S3存储桶不在同一区域,可能导致元数据更新阻塞。确认角色权限,同步资源区域。
  • 优化AWS Wrangler调用参数:
    • 添加use_threads=True启用多线程,加速Parquet写入和元数据操作;
    • 若无需实时同步元数据,可先关闭dataset=True,完成Parquet写入后,单独调用wr.catalog.create_table或wr.catalog.update_table同步元数据,拆分操作避免阻塞。
  • 手动释放内存:在调用to_parquet前,删除不再需要的变量并执行垃圾回收,释放内存资源:
    del df_raw
    import gc
    gc.collect()
    
  • 升级AWS Wrangler版本:部分旧版本的AWS Wrangler在Lambda环境中存在IO阻塞或资源泄漏问题,升级至最新稳定版本(如>=3.0)可修复已知bug。
  • 异步处理元数据同步:如果元数据同步无需实时完成,可将Glue Catalog更新逻辑拆分到另一个Lambda函数,通过SNS或EventBridge触发异步执行,主函数仅完成Parquet写入即可返回,避免超时。

内容的提问来源于stack exchange,提问作者Little Blue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 08:50:02