如何用Python Pandas高效查找DataFrame中的特定数据块?
高效查找Pandas DataFrame中的特定数据块位置
示例数据
原始DataFrame
import pandas as pd df = pd.DataFrame([ [1, 3, 5, 7, 9], [5, 6, 7, 8, 9], [2, 4, 6, 8, 8], [5, 4, 3, 2, 1] ], columns=['A', 'B', 'C', 'D', 'E'])
目标数据块
target = pd.DataFrame([ [7, 8, 9], [6, 8, 8] ])
高效实现方案
循环逐元素匹配效率极低,尤其面对大型DataFrame时。可以利用Numpy向量化操作+滑动窗口视图实现高效匹配,核心是一次性生成所有可能的子矩阵窗口,再批量与目标块比较。
完整代码
import numpy as np import pandas as pd # 初始化数据 df = pd.DataFrame([ [1, 3, 5, 7, 9], [5, 6, 7, 8, 9], [2, 4, 6, 8, 8], [5, 4, 3, 2, 1] ], columns=['A', 'B', 'C', 'D', 'E']) target = pd.DataFrame([ [7, 8, 9], [6, 8, 8] ]) # 转换为Numpy数组,利用底层优化提升速度 arr = df.to_numpy() target_arr = target.to_numpy() target_rows, target_cols = target_arr.shape arr_rows, arr_cols = arr.shape # 用as_strided创建滑动窗口视图(无数据复制,内存占用低) from numpy.lib.stride_tricks import as_strided window_shape = (target_rows, target_cols) strides = arr.strides + arr.strides windows = as_strided( arr, shape=(arr_rows - target_rows + 1, arr_cols - target_cols + 1) + window_shape, strides=strides ) # 批量检查所有窗口与目标块是否完全匹配 matches = (windows == target_arr).all(axis=(2, 3)) # 获取匹配的起始索引 match_indices = np.argwhere(matches) # 转换为DataFrame的行列标签(按需调整) result = [] for i, j in match_indices: result.append({ "起始行索引": df.index[i], "起始列名": df.columns[j], "结束行索引": df.index[i + target_rows - 1], "结束列名": df.columns[j + target_cols - 1] }) # 输出结果 print("匹配到的数据块位置:") for item in result: print(item)
代码说明
- 数据转换:将DataFrame转为Numpy数组,借助Numpy的C级运算优化提升效率。
- 滑动窗口生成:
as_strided仅生成窗口视图,不复制原数据,内存开销极小。 - 批量匹配:通过
all(axis=(2,3))一次性验证窗口内所有元素是否与目标一致,避免Python循环的性能损耗。 - 位置转换:将Numpy索引转为DataFrame的行列标签,贴合业务使用习惯。
输出结果
针对示例数据,运行后会得到:
匹配到的数据块位置: {'起始行索引': 1, '起始列名': 'C', '结束行索引': 2, '结束列名': 'E'}
内容的提问来源于stack exchange,提问作者Kingalione
相关产品推荐
相关产品推荐

