如何使用Pandas写入可被AWS Athena正常读取的Parquet文件
Pandas转Parquet在Athena查询报错问题
转Parquet实现代码
将DataFrame写入Parquet格式字节流的代码如下:
out_buffer = BytesIO() input_datafame.to_parquet(out_buffer, index=False, compression="gzip")
生成文件的Parquet Schema
file schema: schema -------------------------------------------------------------------------------- somID: OPTIONAL INT64 R:0 D:1 SessionID: OPTIONAL INT64 R:0 D:1 JobID: OPTIONAL INT64 R:0 D:1 JobCreationTime: OPTIONAL BINARY L:STRING R:0 D:1 ProcessedId: OPTIONAL BINARY L:STRING R:0 D:1 S3Results: OPTIONAL BINARY L:STRING R:0 D:1
Athena对应表字段定义
`someid` bigint COMMENT '', `sessionid` bigint COMMENT '', `jobid` bigint COMMENT '', `jobcreationtime` string COMMENT '', `processedid` string COMMENT '', `s3results` string COMMENT ''
查询报错信息
执行表查询时抛出如下错误:
HIVE_METASTORE_ERROR: com.amazonaws.services.datacatalog.model.InvalidInputException: Error: type expected at the position 0 of 'integer' but 'integer' is found. (Service: null; Status Code: 0; Error Code: null; Request ID: null; Proxy: null)
初步怀疑方向
目前判断可能的报错原因有两点:
- 整数类型不兼容:Parquet存储的INT64类型与Athena定义的bigint类型存在适配冲突
- 可选类型适配问题:Parquet所有字段均为OPTIONAL可选类型,与Athena表字段的默认属性不匹配
暂未检索到相关已知问题的公开说明,需要确认问题根因。
内容的提问来源于stack exchange,提问作者Istvan
相关产品推荐
相关产品推荐

