如何简洁地道地在Polars DataFrame中循环添加Series匹配行数?
更简洁的Polars实现:循环扩展短Series至DataFrame行数
现有一段Polars代码实现了将长度较短的Series循环填充到DataFrame中,使两者行数匹配,但代码偏冗长。原实现通过行索引取模+表连接完成,具体代码与输出如下:
原代码实现
import polars as pl import numpy as np df = pl.DataFrame(dict( j=np.random.randint(10, 99, 10) )) print('--- df') print(df) s = pl.Series('k', np.random.randint(10, 99, 3)) print('--- s') print(s) dfj = (df .with_row_index() .with_columns( pl.col('index') % len(s) ) .join(s.to_frame().with_row_index(), on='index') .drop('index') ) print('--- dfj') print(dfj)
原代码运行输出
--- df j (i64) 47 22 82 19 85 15 89 74 26 11 shape: (10, 1) --- s shape: (3,) Series: 'k' [i64] [ 86 81 16 ] --- dfj j (i64) k (i64) 47 86 22 81 82 16 19 86 85 81 15 16 89 86 74 81 26 16 11 86 shape: (10, 2)
更简洁的实现方式
方法1:使用repeat_by + cycle(Polars 0.19.0及以上版本支持)
这是最符合Polars惯用风格的写法,直接利用Series的内置方法完成循环扩展:
dfj = df.with_columns(k=s.repeat_by(len(df)).cycle())
说明:repeat_by(len(df))指定要扩展到的总长度,cycle()会自动循环重复原Series的内容,直到达到目标行数。
方法2:通过行索引取模直接索引Series
利用Polars的向量化索引能力,生成循环索引序列直接取值:
dfj = df.with_columns( k=s[pl.int_range(0, len(df)) % len(s)] )
说明:pl.int_range(0, len(df))生成与DataFrame行数一致的连续整数序列,对Series长度取模后得到循环的索引位置,直接用该索引从原Series中取值,实现循环填充。
方法3:手动扩展Series后添加(适合低版本Polars)
如果使用的Polars版本不支持cycle(),可以手动扩展Series到目标长度后再添加:
extended_len = len(df) s_len = len(s) # 计算需要重复的次数,取整后补全剩余部分 extended_s = s * (extended_len // s_len) + s[:extended_len % s_len] dfj = df.with_columns(k=extended_s)
以上几种方法都避免了原代码中with_row_index、join等冗余操作,更简洁且符合Polars的向量化编程风格。
内容的提问来源于stack exchange,提问作者levant pied
相关产品推荐
相关产品推荐

