求助:按指定递推公式批量计算DataFrame各列值(附示例)
Recursive Column Calculation for Pandas DataFrame
The Problem Breakdown
You need to compute updated values for each column in a pandas DataFrame using this chain of recursive formulas:
- For every row:
- New B = A + (1 - A) × Original B
- New C = Updated B + (1 - Updated B) × Original C
- New D = Updated C + (1 - Updated C) × Original D
- This pattern needs to apply to all 11 columns in your actual dataset (the 4-column example is just for testing).
Sample Data
Original Input
| A | B | C | D |
|---|---|---|---|
| 0.2 | 0.4 | 0.8 | 0.5 |
| 0.4 | 0.5 | 0.6 | 0.2 |
| 0.8 | 0.1 | 0.5 | 0.4 |
| 0.3 | 0.4 | 0.1 | 0.8 |
Expected Output
| A | B | C | D |
|---|---|---|---|
| 0.2 | 0.52 | 0.904 | 0.952 |
| 0.4 | 0.7 | 0.88 | 0.904 |
| 0.8 | 0.82 | 0.91 | 0.946 |
| 0.3 | 0.58 | 0.622 | 0.9244 |
Solution Code
Here's a clean, scalable way to implement this logic that works for any number of columns:
import pandas as pd # Initialize your original DataFrame df = pd.DataFrame({ 'A': [0.2, 0.4, 0.8, 0.3], 'B': [0.4, 0.5, 0.1, 0.4], 'C': [0.8, 0.6, 0.5, 0.1], 'D': [0.5, 0.2, 0.4, 0.8] }) # Create a copy to preserve the original data (always a good practice!) df_updated = df.copy() # Loop through columns starting from the second one (since A stays as-is) for col_pos in range(1, df_updated.shape[1]): # Get the already-updated values from the previous column prev_updated_col = df_updated.iloc[:, col_pos - 1] # Get the original values from the current column (don't use the updated copy here!) original_curr_col = df.iloc[:, col_pos] # Apply your recursive formula df_updated.iloc[:, col_pos] = prev_updated_col + (1 - prev_updated_col) * original_curr_col # Print the result to verify print(df_updated)
How This Works
- Data Setup: We start by defining your sample DataFrame. For your 11-column dataset, just add the additional columns to this initial setup.
- Preserve Original Data: Making a copy of the original DataFrame ensures we don't overwrite raw data, which is crucial for debugging or validating results.
- Column Iteration: We loop through each column starting from index 1 (the second column, B). For each column:
- We grab the updated values from the prior column (e.g., when calculating C, we use the already-updated B values)
- We use the original values from the current column (using the original
dfinstead ofdf_updatedprevents accidental overwriting mid-calculation)
- Formula Application: We compute the new values using your specified recursive logic, updating each column in sequence.
Verification
Let's spot-check the first row to confirm:
- Updated B:
0.2 + (1 - 0.2) * 0.4 = 0.2 + 0.32 = 0.52✔️ - Updated C:
0.52 + (1 - 0.52) * 0.8 = 0.52 + 0.384 = 0.904✔️ - Updated D:
0.904 + (1 - 0.904) * 0.5 = 0.904 + 0.048 = 0.952✔️
This matches your expected output exactly, and the code will handle all 11 columns automatically without any extra changes.
内容的提问来源于stack exchange,提问作者bradon Mc
相关产品推荐
相关产品推荐

