Polars如何利用列值作为索引提取列表元素?
问题:基于Polars实现按列值提取列表对应位置元素
测试数据集生成代码
import itertools import numpy as np import polars as pl first = 30 second = 50 third = 40 data = { "a": np.concatenate( (np.repeat(1, first), np.repeat(2, second), np.repeat(3, third)) ), "b": np.concatenate( ( sorted(np.random.randint(1, first, size=first)), sorted(np.random.randint(1, second, size=second)), sorted(np.random.randint(1, third, size=third)), ) ), } d = [ np.tile(np.random.randint(1, first * 2, size=first), (first, 1)).tolist(), np.tile(np.random.randint(1, second * 2, size=second), (second, 1)).tolist(), np.tile(np.random.randint(1, third * 2, size=third), (third, 1)).tolist(), ] data["d"] = list(itertools.chain.from_iterable(d)) df = pl.DataFrame(data) pl_df = df.with_columns([pl.col("a").cum_count().over("a", "b").alias("c")]) pl_df.select(['a', 'b', 'c', "d"]).head()
需求说明
使用Polars 0.20.3版本,需要以b列的值作为1-based索引,提取d列列表中对应位置的元素(例如第一行提取d列表的第23个元素、第二行提取第17个元素),要求不遍历DataFrame的行实现该操作。
解决方案
Polars的list.get方法支持直接传入列作为索引参数,结合Python列表的0-based特性,只需将b列的值减1后传入即可实现需求:
# 提取对应位置元素,新增到extracted列 result_df = pl_df.with_columns( pl.col("d").list.get(pl.col("b") - 1).alias("extracted") ) # 查看结果 result_df.select(['a', 'b', 'c', 'd', 'extracted']).head()
说明
list.get(index):Polars的列表列方法,用于提取列表中指定索引位置的元素pl.col("b") - 1:因为b列的数值是1-based的位置编号,而Python列表是0-based索引,所以需要减1转换为正确的索引值
内容的提问来源于stack exchange,提问作者Maxwell's Daemon
相关产品推荐
相关产品推荐

