如何用Python基于列表变量对DataFrame执行差值计算并新增列?
Solution for Calculating Time-Step Differences in Pandas DataFrame
Got it, let's work through this problem step by step. Here's a straightforward, scalable way to compute those time-based difference columns in your pandas DataFrame:
First, we'll loop through each base variable (OUT and IN), then calculate the difference between consecutive time-suffixed columns (like 3M minus 6M, 6M minus 9M, etc.) and assign these results to new, clearly named columns.
Full Code Implementation
import pandas as pd # Your sample input DataFrame Input_df = pd.DataFrame({ 'ID': ['A', 'B', 'C', 'D'], 'OUT_3M': [2, 3, 2, 3], 'OUT_6M': [3, 3, 3, 3], 'OUT_9M': [4, 5, 6, 7], 'OUT_15M': [6, 7, 6, 7], 'IN_3M': [2, 3, 2, 3], 'IN_6M': [3, 3, 3, 3], 'IN_9M': [4, 5, 6, 7], 'IN_15M': [6, 7, 6, 7] }) # Define your base variables and time suffix order target_vars = ['OUT', 'IN'] time_periods = ['3M', '6M', '9M', '15M'] # Calculate differences for each variable and time pair for var in target_vars: # Iterate over consecutive time periods for idx in range(len(time_periods) - 1): current_period = time_periods[idx] next_period = time_periods[idx + 1] # Create new column name (e.g., OUT_3M-6M) new_col_name = f"{var}_{current_period}-{next_period}" # Compute the difference and add to the DataFrame Input_df[new_col_name] = Input_df[f"{var}_{current_period}"] - Input_df[f"{var}_{next_period}"] # Preview the updated DataFrame print(Input_df)
Breakdown of the Approach
- Define time order: We list the time suffixes in the sequence we need to compare (
3M→6M→9M→15M) to ensure we calculate the correct consecutive differences. - Loop through base variables: For each variable (
OUTandIN), we process all its time-suffixed columns without repeating code. - Dynamic column naming: The new column names follow your requested format (e.g.,
Out_3M-6M) so results are easy to identify. - Efficient calculation: We directly compute and assign the differences to the original DataFrame, avoiding unnecessary copies of data.
When you run this code, your Input_df will include all the new difference columns. For example, the OUT_3M-6M column will show values like -1, 0, -1, 0 for rows A-D respectively, matching the sample data you provided.
内容的提问来源于stack exchange,提问作者Zhoe
相关产品推荐
相关产品推荐

