如何用PyIceberg和Glue Catalog在S3创建Iceberg表?(路径报错)
使用Glue Catalog + PyIceberg在S3创建Iceberg表时遭遇元数据路径空组件错误
我尝试用Glue Catalog和PyIceberg库在S3上创建Iceberg表,定义好schema和分区规则后调用PyIceberg创建表,但多次失败,一直碰到元数据路径空组件相关错误。
我使用的简化代码
import boto3 from pyiceberg.catalog import load_catalog from pyiceberg.schema import Schema from pyiceberg.types import TimestampType, DoubleType, StringType, NestedField from pyiceberg.partitioning import PartitionSpec, PartitionField from pyiceberg.transforms import YearTransform, MonthTransform, DayTransform def create_iceberg_table(): # Replace with your S3 bucket and table names s3_bucket = "my-bucket-name" table_name = "my-table-name" database_name = "iceberg_catalog" # Define the table schema schema = Schema( NestedField(field_id=1, name="field1", field_type=DoubleType(), required=False), NestedField(field_id=2, name="field2", field_type=StringType(), required=False), # ... more fields ... ) # Define the partitioning specification with transformations partition_spec = PartitionSpec( PartitionField(field_id=3, source_id=3, transform=YearTransform(), name="year"), PartitionField(field_id=3, source_id=3, transform=MonthTransform(), name="month"), # ... more partition fields ... ) # Create the Glue client glue_client = boto3.client("glue") # Specify the catalog URI where Glue should store the metadata catalog_uri = f"s3://{s3_bucket}/catalog" # Load the Glue catalog for the specified database catalog = load_catalog("test", client=glue_client, uri=catalog_uri, type="GLUE") # Create the Iceberg table in the Glue Catalog catalog.create_table( identifier=f"{database_name}.{table_name}", schema=schema, partition_spec=partition_spec, location=f"s3://{s3_bucket}/{table_name}/" ) print("Iceberg table created successfully!") if __name__ == "__main__": create_iceberg_table()
报错信息
Traceback (most recent call last): File "/home/workspaceuser/app/create_iceberg_tbl.py", line 72, in <module> create_iceberg_table() File "/home/workspaceuser/app/create_iceberg_tbl.py", line 62, in create_iceberg_table catalog.create_table( File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/catalog/glue.py", line 220, in create_table self._write_metadata(metadata, io, metadata_location) File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/catalog/__init__.py", line 544, in _write_metadata ToOutputFile.table_metadata(metadata, io.new_output(metadata_path)) File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/serializers.py", line 71, in table_metadata with output_file.create(overwrite=overwrite) as output_stream: File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/io/pyarrow.py", line 256, in create if not overwrite and self.exists() is True: File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/io/pyarrow.py", line 200, in exists self._file_info() # raises FileNotFoundError if it does not exist File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/io/pyarrow.py", line 182, in _file_info file_info = self._filesystem.get_file_info(self._path) File "pyarrow/_fs.pyx", line 571, in pyarrow._fs.FileSystem.get_file_info File "pyarrow/error.pxi", line 144, in pyarrow.lib.pyarrow_internal_check_status File "pyarrow/error.pxi", line 100, in pyarrow.lib.check_status pyarrow.lib.ArrowInvalid: Empty path component in path ua-weather-data/hourly_forecasts//metadata/00000-232e3e60-1c1a-4eb8-959e-6940b563acd4.metadata.json
问题根源
报错路径里的hourly_forecasts//metadata出现了连续斜杠,说明路径拼接时产生了空组件,问题出在代码的几个关键配置细节上。
修复步骤
1. 移除location路径末尾的斜杠
你设置的location=f"s3://{s3_bucket}/{table_name}/"末尾多了斜杠,PyIceberg拼接元数据路径时会和metadata目录叠加出//,改成:
location=f"s3://{s3_bucket}/{table_name}"
2. 修正分区字段的field_id冲突
你的PartitionField里多个字段用了同一个field_id=3,违反了Iceberg字段ID唯一性要求,给每个分区字段分配唯一ID:
partition_spec = PartitionSpec( PartitionField(field_id=100, source_id=3, transform=YearTransform(), name="year"), PartitionField(field_id=101, source_id=3, transform=MonthTransform(), name="month"), # 其他分区字段使用不同的field_id )
注意:source_id必须对应schema中实际存在的字段ID,比如你要按时间分区,得先在schema里定义对应字段。
3. 确保Glue Catalog配置正确
Glue Catalog会自动管理元数据存储,不需要手动指定catalog_uri,简化加载代码:
catalog = load_catalog("test", type="GLUE", client=glue_client)
如果需要自定义元数据存储路径,确保路径没有末尾斜杠:
catalog_uri = f"s3://{s3_bucket}/catalog" # 不要加末尾斜杠 catalog = load_catalog("test", type="GLUE", client=glue_client, uri=catalog_uri)
4. 补全schema中缺失的分区源字段
你的schema里没有定义source_id=3对应的字段,分区字段的source_id必须指向已存在的字段,比如添加时间类型字段用于分区:
schema = Schema( NestedField(field_id=1, name="field1", field_type=DoubleType(), required=False), NestedField(field_id=2, name="field2", field_type=StringType(), required=False), NestedField(field_id=3, name="event_time", field_type=TimestampType(), required=False), # 新增分区源字段 )
完整修复后的代码
import boto3 from pyiceberg.catalog import load_catalog from pyiceberg.schema import Schema from pyiceberg.types import TimestampType, DoubleType, StringType, NestedField from pyiceberg.partitioning import PartitionSpec, PartitionField from pyiceberg.transforms import YearTransform, MonthTransform, DayTransform def create_iceberg_table(): # 替换为你的S3桶和表名 s3_bucket = "my-bucket-name" table_name = "my-table-name" database_name = "iceberg_catalog" # 定义表schema,包含分区源字段 schema = Schema( NestedField(field_id=1, name="field1", field_type=DoubleType(), required=False), NestedField(field_id=2, name="field2", field_type=StringType(), required=False), NestedField(field_id=3, name="event_time", field_type=TimestampType(), required=False), ) # 定义分区规则,每个分区字段使用唯一field_id partition_spec = PartitionSpec( PartitionField(field_id=100, source_id=3, transform=YearTransform(), name="year"), PartitionField(field_id=101, source_id=3, transform=MonthTransform(), name="month"), ) # 创建Glue客户端 glue_client = boto3.client("glue") # 加载Glue Catalog catalog = load_catalog("test", type="GLUE", client=glue_client) # 创建Iceberg表,location无末尾斜杠 catalog.create_table( identifier=f"{database_name}.{table_name}", schema=schema, partition_spec=partition_spec, location=f"s3://{s3_bucket}/{table_name}" ) print("Iceberg表创建成功!") if __name__ == "__main__": create_iceberg_table()
内容的提问来源于stack exchange,提问作者Lew
相关产品推荐
相关产品推荐

