Polars:将列表列(含嵌套列表)填充/截断至指定长度
调整Polars DataFrame中所有列表列的长度(截断/用最后一个元素填充)
需求说明
需要将Polars DataFrame中所有列表类型的列(包括嵌套列表列)调整为指定长度:
- 列表长度超过指定值时,截断至该长度
- 列表长度不足时,用列表的最后一个元素重复填充至指定长度
- 非列表类型的列保持不变
解决方案
方法1:使用Python自定义函数(直观易懂)
先定义一个调整列表长度的函数,再对所有列表列应用该函数:
import polars as pl # 创建示例DataFrame df = pl.DataFrame({ "nrs": [[1, 2, 3], [2, 4], [1]], "stuff": [1, 2, 3], "more_stuff": [[[1,1], [2,2], [3,3]], [[4,4], [5,5]], [[6,6]]] }) def adjust_list_length(lst, target_len): if not lst: return lst[:target_len] truncated = lst[:target_len] need = target_len - len(truncated) if need > 0: truncated += [truncated[-1]] * need return truncated # 指定目标长度 N = 2 # 筛选所有列表类型的列 list_cols = [col for col, dtype in df.schema.items() if isinstance(dtype, pl.List)] # 应用调整函数 result_df = df.with_columns( pl.col(col).map_elements(lambda x: adjust_list_length(x, N), return_dtype=pl.col(col).dtype) for col in list_cols ) print(result_df)
方法2:使用Polars向量化表达式(高效,适合大数据量)
利用Polars的内置表达式实现向量化操作,避免逐元素处理的性能损耗:
import polars as pl # 创建示例DataFrame df = pl.DataFrame({ "nrs": [[1, 2, 3], [2, 4], [1]], "stuff": [1, 2, 3], "more_stuff": [[[1,1], [2,2], [3,3]], [[4,4], [5,5]], [[6,6]]] }) # 指定目标长度 N = 2 # 筛选所有列表类型的列 list_cols = [col for col, dtype in df.schema.items() if isinstance(dtype, pl.List)] # 构建调整表达式:生成目标索引,超出列表长度时使用最后一个元素的索引 result_df = df.with_columns( pl.col(col).list.take( pl.int_range(0, N) .zip_with( pl.int_range(0, N) >= pl.col(col).list.len(), pl.col(col).list.len() - 1 ) ).alias(col) for col in list_cols ) print(result_df)
输出结果
两种方法都会得到符合预期的DataFrame:
shape: (3, 3) ┌───────────┬───────┬──────────────────┐ │ nrs ┆ stuff ┆ more_stuff │ │ --- ┆ --- ┆ --- │ │ list[i64] ┆ i64 ┆ list[list[i64]] │ ╞═══════════╪═══════╪══════════════════╡ │ [1, 2] ┆ 1 ┆ [[1, 1], [2, 2]] │ │ [2, 4] ┆ 2 ┆ [[4, 4], [5, 5]] │ │ [1, 1] ┆ 3 ┆ [[6, 6], [6, 6]] │ └───────────┴───────┴──────────────────┘
内容的提问来源于stack exchange,提问作者J.N.
相关产品推荐
相关产品推荐

