如何编写函数实现带位移标记的DataFrame列追加(支持正负位移)
问题描述
现有如下示例Pandas DataFrame代码:
import pandas as pd index_labels = pd.date_range(start='1/1/2018', end='1/08/2018') column_labels = pd.MultiIndex.from_product([['Function A', 'Function B'], ['Fund 1', 'Fund 1', 'Fund 3']]) df = pd.DataFrame(index=index_labels, columns=column_labels) for i in range(len(df.index)): for j in range(len(df.columns)): if df.columns[j][0] == 'Function A': df.iloc[i, j] = i + 1 + j/10 else: df.iloc[i, j] = i + 1 + j/100 df.head()
执行后输出为:
Function A Function B Fund 1 Fund 1 Fund 3 Fund 1 Fund 1 Fund 3 2018-01-01 0.0 0.1 0.2 0.03 0.04 0.05 2018-01-02 1.0 1.1 1.2 1.03 1.04 1.05 2018-01-03 2.0 2.1 2.2 2.03 2.04 2.05 2018-01-04 3.0 3.1 3.2 3.03 3.04 3.05 2018-01-05 4.0 4.1 4.2 4.03 4.04 4.05
需要编写一个名为column_shift的函数,实现以下功能:
- 递归地将后续(或前序,支持负位移)行的列追加到当前行
- 修改列的第二级名称,格式为
Fund ID:位移次数 - 支持正负位移参数
例如调用df2 = column_shift(df, 2)时,预期输出为:
Function A Function B Function A Function B Function A Function B Fund 1 Fund 1 Fund 3 Fund 1 Fund 1 Fund 3 Fund 1:1 Fund 1:1 Fund 3:1 Fund 1:1 Fund 1:1 Fund 3:1 Fund 1:2 Fund 1:2 Fund 3:2 Fund 1:2 Fund 1:2 Fund 3:2 2018-01-01 0.0 0.1 0.2 0.03 0.04 0.05 1.0 1.1 1.2 1.03 1.04 1.05 2.0 2.1 2.2 2.03 2.04 2.05 2018-01-02 1.0 1.1 1.2 1.03 1.04 1.05 2.0 2.1 2.2 2.03 2.04 2.05 3.0 3.1 3.2 3.03 3.04 3.05 2018-01-03 2.0 2.1 2.2 2.03 2.04 2.05 3.0 3.1 3.2 3.03 3.04 3.05 4.0 4.1 4.2 4.03 4.04 4.05
解决方案
以下是实现column_shift函数的代码:
import pandas as pd def column_shift(df, shift_steps): # 初始化待拼接的DataFrame列表,先加入原始数据 dfs_to_concat = [df.copy()] # 根据位移方向确定遍历的步数范围 if shift_steps > 0: step_range = range(1, shift_steps + 1) else: step_range = range(-1, shift_steps - 1, -1) for step in step_range: # 对数据进行行位移 shifted_df = df.shift(periods=step) # 生成新的二级列索引,添加位移标记 new_cols = pd.MultiIndex.from_tuples( [(func, f"{fund}:{step}") for func, fund in shifted_df.columns], names=df.columns.names ) shifted_df.columns = new_cols dfs_to_concat.append(shifted_df) # 按列拼接所有DataFrame result = pd.concat(dfs_to_concat, axis=1) # 移除因位移产生的空值行,保证每行数据完整 if shift_steps > 0: result = result.iloc[:-shift_steps] else: result = result.iloc[-shift_steps:] return result
测试示例
# 生成示例DataFrame index_labels = pd.date_range(start='1/1/2018', end='1/08/2018') column_labels = pd.MultiIndex.from_product([['Function A', 'Function B'], ['Fund 1', 'Fund 1', 'Fund 3']]) df = pd.DataFrame(index=index_labels, columns=column_labels) for i in range(len(df.index)): for j in range(len(df.columns)): if df.columns[j][0] == 'Function A': df.iloc[i, j] = i + 1 + j/10 else: df.iloc[i, j] = i + 1 + j/100 # 测试正位移(2步) df2 = column_shift(df, 2) print("正位移2步结果:") print(df2) # 测试负位移(-1步) df3 = column_shift(df, -1) print("\n负位移1步结果:") print(df3)
函数说明
- 参数:
df:输入的带有二级列索引的Pandas DataFrameshift_steps:位移步数,正数表示取后续行,负数表示取前序行
- 核心逻辑:
- 复制原始数据作为基础拼接项
- 根据位移方向遍历每一步,生成位移后的DataFrame并修改列索引
- 拼接所有位移后的结果
- 移除包含空值的不完整行,确保输出数据的有效性
内容的提问来源于stack exchange,提问作者John
相关产品推荐
相关产品推荐

