You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何简洁地道地在Polars DataFrame中循环添加Series匹配行数?

更简洁的Polars实现:循环扩展短Series至DataFrame行数

现有一段Polars代码实现了将长度较短的Series循环填充到DataFrame中,使两者行数匹配,但代码偏冗长。原实现通过行索引取模+表连接完成,具体代码与输出如下:

原代码实现

import polars as pl
import numpy as np

df = pl.DataFrame(dict(
  j=np.random.randint(10, 99, 10)
))
print('--- df')
print(df)

s = pl.Series('k', np.random.randint(10, 99, 3))
print('--- s')
print(s)

dfj = (df
  .with_row_index()
  .with_columns(
    pl.col('index') % len(s)
  )
  .join(s.to_frame().with_row_index(), on='index')
  .drop('index')
)
print('--- dfj')
print(dfj)

原代码运行输出

--- df
 j (i64)
 47
 22
 82
 19
 85
 15
 89
 74
 26
 11
shape: (10, 1)
--- s
shape: (3,)
Series: 'k' [i64]
[
        86
        81
        16
]
--- dfj
 j (i64)  k (i64)
 47       86
 22       81
 82       16
 19       86
 85       81
 15       16
 89       86
 74       81
 26       16
 11       86
shape: (10, 2)

更简洁的实现方式

方法1:使用repeat_by + cycle(Polars 0.19.0及以上版本支持)

这是最符合Polars惯用风格的写法,直接利用Series的内置方法完成循环扩展:

dfj = df.with_columns(k=s.repeat_by(len(df)).cycle())

说明:repeat_by(len(df))指定要扩展到的总长度,cycle()会自动循环重复原Series的内容,直到达到目标行数。

方法2:通过行索引取模直接索引Series

利用Polars的向量化索引能力,生成循环索引序列直接取值:

dfj = df.with_columns(
    k=s[pl.int_range(0, len(df)) % len(s)]
)

说明:pl.int_range(0, len(df))生成与DataFrame行数一致的连续整数序列,对Series长度取模后得到循环的索引位置,直接用该索引从原Series中取值,实现循环填充。

方法3:手动扩展Series后添加(适合低版本Polars)

如果使用的Polars版本不支持cycle(),可以手动扩展Series到目标长度后再添加:

extended_len = len(df)
s_len = len(s)
# 计算需要重复的次数,取整后补全剩余部分
extended_s = s * (extended_len // s_len) + s[:extended_len % s_len]
dfj = df.with_columns(k=extended_s)

以上几种方法都避免了原代码中with_row_index、join等冗余操作,更简洁且符合Polars的向量化编程风格。

内容的提问来源于stack exchange,提问作者levant pied

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 05:13:21