求助:为DataFrame多列新增行差计算列(row[n+1]-row[n])的实现
Hey there, this is a super common task when working with sequential financial data (like your stock earnings use case), and pandas has a built-in tool that makes this straightforward. Let me walk you through two practical approaches to get exactly what you need—adding 6 new columns following the stock_name + "_Earning" rule, while keeping all your original data intact.
Approach 1: Loop Through Columns (Great for Step-by-Step Clarity)
If you prefer explicit, easy-to-follow code, looping through each column lets you control every part of the process:
Step-by-Step Code
First, let's create a sample DataFrame to simulate your stock data:
import pandas as pd import numpy as np # 模拟6只股票的原始数据(10行示例) np.random.seed(42) df = pd.DataFrame({ 'StockA': np.random.randint(100, 200, 10), 'StockB': np.random.randint(80, 180, 10), 'StockC': np.random.randint(120, 220, 10), 'StockD': np.random.randint(90, 190, 10), 'StockE': np.random.randint(110, 210, 10), 'StockF': np.random.randint(70, 170, 10) })
Now calculate the row-to-row differences and add the new columns:
# 遍历每一列,计算当前行与前一行的差值 for col in df.columns: # 按照规则命名新列 new_col_name = f"{col}_Earning" # 使用diff():默认就是计算row[n+1] - row[n]的差值 df[new_col_name] = df[col].diff()
Approach 2: Batch Processing (Cleaner & More Efficient)
For a concise, one-liner-style solution, you can generate all difference columns at once and merge them with the original DataFrame:
# 批量生成所有差值列,并按规则重命名 earning_columns = df.diff().rename(columns=lambda x: f"{x}_Earning") # 合并原数据和新列(axis=1表示按列拼接) df = pd.concat([df, earning_columns], axis=1)
Key Notes
- NaN Handling: The first row of each new
_Earningcolumn will beNaN(since there's no previous row to compare). If you need to fill this gap (e.g., with 0), just add.fillna(0)to thediff()call:df[new_col_name] = df[col].diff().fillna(0) - Preserve Original Data: Both methods keep all your original columns intact—they only append the new difference columns to the end of the DataFrame.
- Naming Accuracy: Using f-strings (Python 3.6+) or lambda functions ensures your new columns strictly follow the
stock_name + "_Earning"convention.
Verify the Result
To confirm all 6 new columns were added, print the DataFrame's column list:
print(df.columns) # Output: Index(['StockA', 'StockB', 'StockC', 'StockD', 'StockE', 'StockF', # 'StockA_Earning', 'StockB_Earning', 'StockC_Earning', # 'StockD_Earning', 'StockE_Earning', 'StockF_Earning'], # dtype='object')
内容的提问来源于stack exchange,提问作者3kstc

