You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PyIceberg和Glue Catalog在S3创建Iceberg表?(路径报错)

使用Glue Catalog + PyIceberg在S3创建Iceberg表时遭遇元数据路径空组件错误

我尝试用Glue Catalog和PyIceberg库在S3上创建Iceberg表,定义好schema和分区规则后调用PyIceberg创建表,但多次失败,一直碰到元数据路径空组件相关错误。

我使用的简化代码

import boto3
from pyiceberg.catalog import load_catalog
from pyiceberg.schema import Schema
from pyiceberg.types import TimestampType, DoubleType, StringType, NestedField
from pyiceberg.partitioning import PartitionSpec, PartitionField
from pyiceberg.transforms import YearTransform, MonthTransform, DayTransform

def create_iceberg_table():
    # Replace with your S3 bucket and table names
    s3_bucket = "my-bucket-name"
    table_name = "my-table-name"
    database_name = "iceberg_catalog"

    # Define the table schema
    schema = Schema(
        NestedField(field_id=1, name="field1", field_type=DoubleType(), required=False),
        NestedField(field_id=2, name="field2", field_type=StringType(), required=False),
        # ... more fields ...
    )

    # Define the partitioning specification with transformations
    partition_spec = PartitionSpec(
        PartitionField(field_id=3, source_id=3, transform=YearTransform(), name="year"),
        PartitionField(field_id=3, source_id=3, transform=MonthTransform(), name="month"),
        # ... more partition fields ...
    )

    # Create the Glue client
    glue_client = boto3.client("glue")

    # Specify the catalog URI where Glue should store the metadata
    catalog_uri = f"s3://{s3_bucket}/catalog"
    # Load the Glue catalog for the specified database
    catalog = load_catalog("test", client=glue_client, uri=catalog_uri, type="GLUE")

    # Create the Iceberg table in the Glue Catalog
    catalog.create_table(
        identifier=f"{database_name}.{table_name}",
        schema=schema,
        partition_spec=partition_spec,
        location=f"s3://{s3_bucket}/{table_name}/"
    )

    print("Iceberg table created successfully!")

if __name__ == "__main__":
    create_iceberg_table()

报错信息

Traceback (most recent call last):
  File "/home/workspaceuser/app/create_iceberg_tbl.py", line 72, in <module>
    create_iceberg_table()
  File "/home/workspaceuser/app/create_iceberg_tbl.py", line 62, in create_iceberg_table
    catalog.create_table(
  File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/catalog/glue.py", line 220, in create_table
    self._write_metadata(metadata, io, metadata_location)
  File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/catalog/__init__.py", line 544, in _write_metadata
    ToOutputFile.table_metadata(metadata, io.new_output(metadata_path))
  File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/serializers.py", line 71, in table_metadata
    with output_file.create(overwrite=overwrite) as output_stream:
  File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/io/pyarrow.py", line 256, in create
    if not overwrite and self.exists() is True:
  File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/io/pyarrow.py", line 200, in exists
    self._file_info()  # raises FileNotFoundError if it does not exist
  File "/home/workspaceuser/layers/paketo-buildpacks_cpython/cpython/lib/python3.8/site-packages/pyiceberg/io/pyarrow.py", line 182, in _file_info
    file_info = self._filesystem.get_file_info(self._path)
  File "pyarrow/_fs.pyx", line 571, in pyarrow._fs.FileSystem.get_file_info
  File "pyarrow/error.pxi", line 144, in pyarrow.lib.pyarrow_internal_check_status
  File "pyarrow/error.pxi", line 100, in pyarrow.lib.check_status
pyarrow.lib.ArrowInvalid: Empty path component in path ua-weather-data/hourly_forecasts//metadata/00000-232e3e60-1c1a-4eb8-959e-6940b563acd4.metadata.json

问题根源

报错路径里的hourly_forecasts//metadata出现了连续斜杠,说明路径拼接时产生了空组件,问题出在代码的几个关键配置细节上。

修复步骤

1. 移除location路径末尾的斜杠

你设置的location=f"s3://{s3_bucket}/{table_name}/"末尾多了斜杠,PyIceberg拼接元数据路径时会和metadata目录叠加出//,改成:

location=f"s3://{s3_bucket}/{table_name}"

2. 修正分区字段的field_id冲突

你的PartitionField里多个字段用了同一个field_id=3,违反了Iceberg字段ID唯一性要求,给每个分区字段分配唯一ID:

partition_spec = PartitionSpec(
    PartitionField(field_id=100, source_id=3, transform=YearTransform(), name="year"),
    PartitionField(field_id=101, source_id=3, transform=MonthTransform(), name="month"),
    # 其他分区字段使用不同的field_id
)

注意:source_id必须对应schema中实际存在的字段ID,比如你要按时间分区,得先在schema里定义对应字段。

3. 确保Glue Catalog配置正确

Glue Catalog会自动管理元数据存储,不需要手动指定catalog_uri,简化加载代码:

catalog = load_catalog("test", type="GLUE", client=glue_client)

如果需要自定义元数据存储路径,确保路径没有末尾斜杠:

catalog_uri = f"s3://{s3_bucket}/catalog"  # 不要加末尾斜杠
catalog = load_catalog("test", type="GLUE", client=glue_client, uri=catalog_uri)

4. 补全schema中缺失的分区源字段

你的schema里没有定义source_id=3对应的字段,分区字段的source_id必须指向已存在的字段,比如添加时间类型字段用于分区:

schema = Schema(
    NestedField(field_id=1, name="field1", field_type=DoubleType(), required=False),
    NestedField(field_id=2, name="field2", field_type=StringType(), required=False),
    NestedField(field_id=3, name="event_time", field_type=TimestampType(), required=False),  # 新增分区源字段
)

完整修复后的代码

import boto3
from pyiceberg.catalog import load_catalog
from pyiceberg.schema import Schema
from pyiceberg.types import TimestampType, DoubleType, StringType, NestedField
from pyiceberg.partitioning import PartitionSpec, PartitionField
from pyiceberg.transforms import YearTransform, MonthTransform, DayTransform

def create_iceberg_table():
    # 替换为你的S3桶和表名
    s3_bucket = "my-bucket-name"
    table_name = "my-table-name"
    database_name = "iceberg_catalog"

    # 定义表schema,包含分区源字段
    schema = Schema(
        NestedField(field_id=1, name="field1", field_type=DoubleType(), required=False),
        NestedField(field_id=2, name="field2", field_type=StringType(), required=False),
        NestedField(field_id=3, name="event_time", field_type=TimestampType(), required=False),
    )

    # 定义分区规则,每个分区字段使用唯一field_id
    partition_spec = PartitionSpec(
        PartitionField(field_id=100, source_id=3, transform=YearTransform(), name="year"),
        PartitionField(field_id=101, source_id=3, transform=MonthTransform(), name="month"),
    )

    # 创建Glue客户端
    glue_client = boto3.client("glue")

    # 加载Glue Catalog
    catalog = load_catalog("test", type="GLUE", client=glue_client)

    # 创建Iceberg表,location无末尾斜杠
    catalog.create_table(
        identifier=f"{database_name}.{table_name}",
        schema=schema,
        partition_spec=partition_spec,
        location=f"s3://{s3_bucket}/{table_name}"
    )

    print("Iceberg表创建成功!")

if __name__ == "__main__":
    create_iceberg_table()

内容的提问来源于stack exchange,提问作者Lew

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 16:55:15