使用PyIceberg写入Map类型数据时遭遇Schema校验错误的问题咨询
PyIceberg写入Map类型数据时遭遇Schema校验错误的问题咨询
我不确定这是个bug还是我数据结构没写对,找了一圈也没找到Map类型写入的示例。
我先创建了一个带Map字段的简单表,代码如下:
from pyiceberg.schema import Schema from pyiceberg.types import StringType, MapType, NestedField map_type = MapType(key_id=1001, key_type=StringType(), value_id=1002, value_type=StringType()) schema = Schema(NestedField(field_id=1, name='my_map', field_type=map_type)) table = catalog.create_table(..., schema=schema)
创建后的表结构如下:
table map.test( 1: my_map: optional map<string, string> ), partition by: [], sort order: [], snapshot: null
之后我用转换后的Schema构造了一个Arrow表:
data = {'my_map': [{'symbol': 'BTC'}]} pa_table = pa.Table.from_pydict(data, schema=schema.as_arrow())
这个Arrow表的结构是:
pyarrow.Table my_map: map<large_string, large_string> child 0, entries: struct<key: large_string not null, value: large_string not null> not null child 0, key: large_string not null child 1, value: large_string not null ---- my_map: [[keys:["symbol"]values:["BTC"]]]
但当我尝试写入数据时,Schema校验报错说我没有提供key和value字段:
>>> table.append(pa_table) ┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓ ┃ ┃ Table field ┃ Dataframe field ┃ ┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┩ │ ✅ │ 1: my_map: optional map<string, string> │ 1: my_map: optional map<string, string> │ │ ❌ │ 2: key: required string │ Missing │ │ ❌ │ 3: value: required string │ Missing │ └────┴─────────────────────────────────────────┴─────────────────────────────────────────┘
备注:内容来源于stack exchange,提问作者bphi
相关产品推荐
相关产品推荐

