如何在Python中高效生成滚动子序列的DataFrame
高效生成滑动窗口序列DataFrame的方法
针对你用循环+iloc处理大数据量效率极低的问题,推荐使用numpy的内存视图技巧(as_strided)来实现,这种方法无需复制数据,速度能提升几个数量级,处理十万级数据也能瞬间完成。
核心实现代码
import pandas as pd import numpy as np from numpy.lib.stride_tricks import as_strided # 模拟你的原始DataFrame df = pd.DataFrame({'col0': range(1, 1001)}) def create_sliding_window_df(arr, window_length): # 获取数组的内存步长(字节单位) elem_stride = arr.strides[0] # 计算滑动窗口的总行数:总长度 - 窗口长度 + 1 total_rows = arr.size - window_length + 1 # 创建滑动窗口的内存视图(无数据复制) window_array = as_strided(arr, shape=(total_rows, window_length), strides=(elem_stride, elem_stride)) # 转换为DataFrame并设置列名 return pd.DataFrame(window_array, columns=[f'col{i}' for i in range(window_length)]) # 对应你的需求:生成5行、996列的结果,所以窗口长度设为996 result_df = create_sliding_window_df(df['col0'].values, window_length=996)
为什么这个方法高效?
as_strided直接基于原始数组的内存创建视图,不复制任何数据,时间复杂度为O(N);- 循环+
iloc的方法需要反复切片、创建新对象,时间复杂度为O(N*K),数据量越大效率差距越明显。
备选方案(pandas shift方法)
如果不想用numpy,也可以用pandas的shift拼接列,但效率略低于numpy方法,适合窗口长度较小的场景:
window_length = 996 # 生成window_length列,每列是原数据依次偏移i位 result_df = pd.concat([df['col0'].shift(i) for i in range(window_length)], axis=1) # 去掉包含NaN的行(滑动窗口的无效部分)并重置索引 result_df = result_df.dropna().reset_index(drop=True) # 设置列名 result_df.columns = [f'col{i}' for i in range(window_length)]
内容的提问来源于stack exchange,提问作者ntintel
相关产品推荐
相关产品推荐

