如何将含嵌套字典的列结构转换为正常的Polars DataFrame?
如何将含嵌套字典的列结构转换为正常的Polars DataFrame?
这个问题我之前也碰到过,本质是Polars默认会把内层字典识别成单个struct类型元素,而不是对应多行的序列。咱们可以用两种简单的方式解决:
方法1:将每个嵌套字典转为Series后拼接
这是最直接的处理方式,遍历外层字典的每个键值对,把内层字典转换成Polars的Series,再组合成DataFrame:
import polars as pl d = {'col1': {'0':'A','1':'B','2':'C'}, 'col2': {'0':1,'1':2,'2':3}} # 遍历列名和对应的嵌套字典,转为Series后构建DataFrame df = pl.DataFrame({col_name: pl.Series(col_data) for col_name, col_data in d.items()}) print(df)
执行后就能得到你想要的正常DataFrame:
shape: (3, 2) ┌──────┬──────┐ │ col1 ┆ col2 │ │ --- ┆ --- │ │ str ┆ i64 │ ╞══════╪══════╡ │ A ┆ 1 │ │ B ┆ 2 │ │ C ┆ 3 │ └──────┴──────┘
方法2:重构字典结构后用orient='index'转换
如果想换一种思路,可以先把原字典重构为以索引为键、列名-值为内层的结构,再用from_dict的orient='index'参数生成DataFrame:
import polars as pl d = {'col1': {'0':'A','1':'B','2':'C'}, 'col2': {'0':1,'1':2,'2':3}} # 重构字典:键是原索引,值是各列对应的元素 indexed_dict = { idx: {col: d[col][idx] for col in d} for idx in d['col1'] # 假设所有列的索引一致 } # 用orient='index'按索引行来构建DataFrame df = pl.DataFrame.from_dict(indexed_dict, orient='index') print(df)
这个方法也能得到完全一样的结果,适合需要先处理索引逻辑的场景。
为什么原方法不行?
Polars的pl.DataFrame(d)或pl.from_dict(d)默认会把每个内层字典当作单个结构化数据(也就是struct类型),所以最终每列只会有一个包含所有值的struct元素,而不是展开成多行。通过显式转换为Series,我们相当于告诉Polars:这个内层字典是一列的多个行值,需要对应到不同的行。
备注:内容来源于stack exchange,提问作者Ghost
相关产品推荐
相关产品推荐

