Polars与Pandas中pivot函数不支持List[str]类型的问题求助
解决Polars/Pandas中Pivot因List类型索引报错的问题
问题背景
给定数据:
data = {"error":[["x","z"],None,["x","z"],None], "X" : ["x","p","x","p"], "Y" : ["y","q","y","q"], "Z": ["z","r","z","r"], "time": ["Jan","Jan","Feb","Feb"], "value": [10,20,15,19]}
对应数据表结构:
| error | X | Y | Z | time | value |
|---|---|---|---|---|---|
| ["x","z"] | x | y | z | Jan | 10 |
| None | p | q | r | Jan | 20 |
| ["x","z"] | x | y | z | Feb | 15 |
| None | p | q | r | Feb | 19 |
执行Polars的pivot代码时:
import polars as pl df = pl.DataFrame(data) df.pivot(values="value",columns="time",index=["X","Y","Z","error"])
会抛出错误提示不支持List[str]类型;在Pandas中执行类似操作也会失败,核心原因是列表(List)属于不可哈希类型,而pivot操作的索引列必须是可哈希的。
解决方案
将error列的列表类型转换为可哈希类型(如元组、JSON字符串)即可解决问题,以下分Polars和Pandas两种场景给出实现:
Polars 实现
方案1:将列表转为元组
元组是可哈希类型,直接转换后即可正常pivot:
import polars as pl data = {"error":[["x","z"],None,["x","z"],None], "X" : ["x","p","x","p"], "Y" : ["y","q","y","q"], "Z": ["z","r","z","r"], "time": ["Jan","Jan","Feb","Feb"], "value": [10,20,15,19]} df = pl.DataFrame(data) # 转换error列:列表转元组,None保持不变 df_processed = df.with_columns( pl.col("error").map_elements(lambda x: tuple(x) if x is not None else x, return_dtype=pl.Object) ) # 执行pivot result = df_processed.pivot(values="value", columns="time", index=["X","Y","Z","error"]) print(result)
方案2:将列表转为JSON字符串(推荐,格式更直观)
用JSON字符串序列化列表,既保证可哈希,又保留原列表的可读格式:
import polars as pl import json df_processed = df.with_columns( pl.col("error").map_elements( lambda x: json.dumps(x) if x is not None else x, return_dtype=pl.String ) ) result = df_processed.pivot(values="value", columns="time", index=["X","Y","Z","error"]) print(result)
执行后输出结果:
shape: (2, 6) X Y Z error Jan Feb str str str str i64 i64 --- --- --- ----------- --- --- x y z ["x","z"] 10 15 p q r null 20 19
Pandas 实现
方案1:将列表转为元组
import pandas as pd data = {"error":[["x","z"],None,["x","z"],None], "X" : ["x","p","x","p"], "Y" : ["y","q","y","q"], "Z": ["z","r","z","r"], "time": ["Jan","Jan","Feb","Feb"], "value": [10,20,15,19]} df = pd.DataFrame(data) # 转换error列 df["error"] = df["error"].apply(lambda x: tuple(x) if x is not None else x) # 执行pivot result = df.pivot(values="value", columns="time", index=["X","Y","Z","error"]) print(result)
方案2:将列表转为JSON字符串
import pandas as pd import json df["error"] = df["error"].apply(lambda x: json.dumps(x) if x is not None else x) result = df.pivot(values="value", columns="time", index=["X","Y","Z","error"]) print(result)
执行后输出结果:
Jan Feb X Y Z error x y z ("x", "z") 10 15 p q r None 20 19
内容的提问来源于stack exchange,提问作者Alby
相关产品推荐
相关产品推荐

