如何基于(start,length)元组列表提取Polars DataFrame多段切片?
从Polars DataFrame提取指定多段切片
方法一:利用slice和concat拼接片段
直接遍历切片元组列表,用slice提取每段数据,再通过concat合并结果:
import polars as pl df = pl.DataFrame(data={"col1": range(10)}) slices = [(1,2), (5,3)] result = pl.concat([df.slice(start, length) for start, length in slices]) print(result)
运行后输出目标结果:
┌──────┐ │ col1 │ │ --- │ │ i64 │ ╞══════╡ │ 1 │ │ 2 │ │ 5 │ │ 6 │ │ 7 │ └──────┘
方法二:通过行索引筛选
先计算所有需要提取的行索引,再直接按索引选取行:
import polars as pl df = pl.DataFrame(data={"col1": range(10)}) slices = [(1,2), (5,3)] # 生成所有目标行的索引 target_indices = [] for start, length in slices: target_indices.extend(range(start, start + length)) result = df[target_indices] print(result)
两种方法都能得到预期结果:第一种更贴合题目中slice的参数规则,代码更简洁;第二种适合需要对索引做额外处理的场景。
内容的提问来源于stack exchange,提问作者Andi
相关产品推荐
相关产品推荐

