Polars DataFrame如何按pl.List数据类型筛选列?
如何筛选Polars DataFrame中类型为pl.List的列
直接使用df.select(pl.col(pl.List))无法筛选List类型列,这是因为Polars的List是参数化泛型类型(比如示例中的list[str]),pl.List作为基类,无法直接匹配到具体的带内部类型的List列。
以下是两种可行的解决方案:
方法1:通过filter结合类型判断筛选列
利用pl.col().filter()配合dtype.is_a(pl.List)来识别所有List类型的列:
import polars as pl df = pl.DataFrame({"foo": [[c] for c in ["100CT pen", "pencils 250CT", "what 125CT soever", "this is a thing"]]} ) # 筛选List类型列 result = df.select(pl.col().filter(lambda col: col.dtype.is_a(pl.List))) print(result)
输出:
shape: (4, 1) ┌─────────────────────────┐ │ foo │ │ --- │ │ list[str] │ ╞═════════════════════════╡ │ ["100CT pen"] │ │ ["pencils 250CT"] │ │ ["what 125CT soever"] │ │ ["this is a thing"] │ └─────────────────────────┘
方法2:遍历Schema筛选列名
先遍历DataFrame的schema,提取所有类型为List的列名,再通过列名选择:
# 获取所有List类型的列名 list_cols = [col_name for col_name, dtype in df.schema.items() if dtype.is_a(pl.List)] # 选择列 result = df.select(list_cols) print(result)
输出和方法1一致。
原理说明
Polars中具体的List类型(如list[str]、list[i64])都是pl.List的子类,dtype.is_a(pl.List)会检查当前类型是否继承自pl.List,因此能准确匹配所有List类型的列,而直接使用pl.col(pl.List)只会匹配未参数化的pl.List类型(实际中几乎不会用到),所以无法生效。
内容的提问来源于stack exchange,提问作者Björn
相关产品推荐
相关产品推荐

