You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将JSON转为Parquet时遇'Glue Table不存在'错误求助

问题:JSON转Parquet时触发“Glue Table does not exist in catalog”错误

尝试将S3中的JSON文件转换为Parquet格式,使用AWS Lambda结合awswrangler库,运行时遇到错误提示“Glue Table does not exist in catalog”。已在Lambda配置中添加环境变量,权限策略包含AmazonS3FullAccess和AmazonGlueServiceRole,期望在指定新Bucket中生成Parquet文件。

代码实现

import awswrangler as wr
import pandas as pd
import urllib.parse
import os

os_input_s3_file_path = os.environ['s3_cleaned_layer']
os_input_glue_catalog_db_name = os.environ['glue_catalog_db_name']
os_input_glue_catalog_table_name = os.environ['glue_catalog_table_name']
os_input_write_data_operation = os.environ['write_data_operation']

def lambda_handler(event, context):
    
    #Get object and key name
    bucket = event['Records'][0]['s3']['bucket']['name']
    key= urllib.parse.unquote_plus(event['Records'][0]['s3']['object']['key'], encoding = 'utf-8')             
    
    try:
        
        #creating df from content
        df_raw = wr.s3.read_json('s3://{}/{}'.format(bucket, key))
        
        #get item from raw df
        df_step_1 = pd.json_normalize(df_raw['items'])
        
        #write to s3
        wr_response = wr.s3.to_parquet(
            df=df_step_1,
            dataset=True,
            database=os_input_glue_catalog_db_name,
            table=os_input_glue_catalog_table_name,
            mode=os_input_write_data_operation
            )
        
        return wr_response
        
    except Exception as e:
        print(e)
        print('Error getting object {}  {}'.format(key, bucket))
        raise e

解决方案

  • 确认Glue Catalog中数据库和表的存在性
    登录AWS Glue控制台,检查环境变量glue_catalog_db_name对应的数据库是否存在,且glue_catalog_table_name对应的表已在该数据库下创建。如果表未创建,可通过两种方式处理:

    1. 手动在Glue控制台创建表,确保表结构与DataFramedf_step_1的字段匹配;
    2. 修改wr.s3.to_parquet调用,添加path和table_type参数,让awswrangler自动创建外部表:
      wr_response = wr.s3.to_parquet(
          df=df_step_1,
          dataset=True,
          database=os_input_glue_catalog_db_name,
          table=os_input_glue_catalog_table_name,
          mode=os_input_write_data_operation,
          path=os_input_s3_file_path,
          table_type="EXTERNAL_TABLE"
      )
      
  • 验证环境变量的准确性
    检查Lambda环境变量中glue_catalog_db_name和glue_catalog_table_name的拼写、大小写是否与Glue控制台中的完全一致,避免因名称不匹配导致的表找不到问题。

  • 检查Lambda执行角色的权限
    虽然已附加AmazonGlueServiceRole,但需确认该角色包含Glue Catalog的必要操作权限(如glue:CreateTable、glue:GetTable、glue:UpdateTable等)。可临时添加AWSGlueFullAccess策略排查权限问题,后续再按需缩小权限范围。

  • 确认S3路径与区域配置
    确保os_input_s3_file_path指向目标Bucket的正确路径,且该Bucket与Lambda函数处于同一AWS区域(若跨区域,需在wr.s3.to_parquet中指定region参数)。


内容的提问来源于stack exchange,提问作者Mayur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 17:55:20