用Python函数高效实现通用DataFrame的指定规则新行添加
为任意尺寸DataFrame添加符合特定规则的新行
需求说明
给定n行m列的DataFrame,需在末尾新增一行,新行每个单元格的取值规则为:第k列(从1开始计数)的单元格,取原DataFrame第k列中比最后一行早k行的值(即该列的倒数第k+1个元素)。
以6×5的DataFrame为例:
原始DataFrame
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30
期望结果(替换原最后一行的效果,实际可选择新增或替换)
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 21 17 13 9 5
高效解决方案
以下函数基于pandas向量化操作实现,避免循环,保证最优运行效率,且支持任意行列数的DataFrame:
1. 新增行版本
import pandas as pd def add_special_row(df): n_rows, n_cols = df.shape # 计算每个列对应的目标行索引:0-based列索引j对应行索引为n_rows - j - 2 target_rows = [n_rows - j - 2 for j in range(n_cols)] # 一次性提取所有目标位置的元素组成新行 new_row = df.lookup(target_rows, df.columns) # 拼接新行到原DataFrame末尾并重置索引 return pd.concat([df, pd.DataFrame([new_row], columns=df.columns)], ignore_index=True)
2. 替换最后一行版本
若需求是替换原最后一行而非新增,使用以下函数:
def replace_last_row(df): n_rows, n_cols = df.shape target_rows = [n_rows - j - 2 for j in range(n_cols)] new_row = df.lookup(target_rows, df.columns) df.iloc[-1] = new_row return df
代码说明
- 效率优势:使用
df.lookup()进行向量化索引,这是pandas中针对多位置元素提取的高效方法,性能远优于逐列循环,尤其适合大尺寸DataFrame。 - 通用性:函数自动获取输入DataFrame的行列数,无需手动传入参数,适配任意规格的DataFrame。
- 逻辑对应:对于0-based列索引
j,目标行索引为n_rows - j - 2,等价于1-based列序号k=j+1时,取原列倒数第k+1个元素,完全匹配需求规则。
测试示例
# 构造示例DataFrame sample_data = [ [1,2,3,4,5], [6,7,8,9,10], [11,12,13,14,15], [16,17,18,19,20], [21,22,23,24,25], [26,27,28,29,30] ] sample_df = pd.DataFrame(sample_data) # 测试新增行 result_add = add_special_row(sample_df) print("新增行结果:") print(result_add) # 测试替换最后一行 result_replace = replace_last_row(sample_df.copy()) print("\n替换最后一行结果:") print(result_replace)
内容的提问来源于stack exchange,提问作者Rebel
相关产品推荐
相关产品推荐

