如何在Polars中根据另一列的值动态获取对应列内容?
解决方案:根据指定列名动态提取对应计数列的值
要实现从largest列指定的列名对应的*_count列提取值,且避免逐个列判断的繁琐操作,可利用Polars的结构体(Struct)特性进行向量化处理,高效适配列数较多的场景:
代码实现
import polars as pl # 初始化输入DataFrame df = pl.DataFrame({ "foo": [1, 2, 3, 4], "foo_count": [23, 45, 234, 55], "bar": [4, 6, 9, 2], "bar_count": [43, 45, 453, 67], "baz": [5, 1, 15, 3], "baz_count": [64, 43, 231, 94], "largest": ["baz", "bar", "baz", "foo"] }) # 动态提取对应计数列的值 result = df.with_columns( # 将所有*_count列打包为结构体 pl.struct(pl.col("*_count")) # 根据largest列的值提取结构体中对应键的内容 .get(pl.col("largest")) .alias("largest_count") ) print(result)
代码说明
- 打包结构体:
pl.struct(pl.col("*_count"))自动匹配所有后缀为_count的列,将它们打包成每行对应的结构体对象,结构体的键即为原列名(如foo_count、bar_count)。 - 动态取值:
.get(pl.col("largest"))逐行读取largest列的内容(如baz、bar),并以此为键从结构体中取出对应的值,实现按需动态提取。 - 命名新列:通过
.alias("largest_count")将提取结果命名为目标列。
输出结果
运行代码后得到的DataFrame与需求一致:
| foo | foo_count | bar | bar_count | baz | baz_count | largest | largest_count |
|---|---|---|---|---|---|---|---|
| 1 | 23 | 4 | 43 | 5 | 64 | baz | 64 |
| 2 | 45 | 6 | 45 | 1 | 43 | bar | 45 |
| 3 | 234 | 9 | 453 | 15 | 231 | baz | 231 |
| 4 | 55 | 2 | 67 | 3 | 94 | foo | 55 |
(注:原需求中第四行largest_count标注为4应为笔误,实际对应foo_count的值是55)
内容的提问来源于stack exchange,提问作者Kazdegotepu
相关产品推荐
相关产品推荐

