如何为DataFrame按A分组填充B=Y对应的D值至新列E?
Got it, let's tackle this problem. The goal is to populate a new column E where each row gets the D value corresponding to B='Y' within the same A group. Here are two straightforward ways to do this:
Method 1: Using groupby + transform
This approach works by grouping the DataFrame by column A, then extracting the D value for B='Y' in each group and broadcasting it to all rows in that group.
First, let's set up the sample DataFrame to test with:
import pandas as pd data = { 'A': ['2002-01-13 15:00:00']*4 + ['2002-01-14 16:00:00']*4, 'B': ['X', 'Y', 'X', 'X', 'X', 'Y', 'X', 'X'], 'C': [110, 120, 130, 140, 110, 120, 130, 140], 'D': [3.9, 1.9, 8.0, 9.0, 0.2, 7.0, 1.6, 3.4] } df = pd.DataFrame(data)
Then run the core code to create column E:
# For each group in 'A', grab the D value where B='Y' and apply it to all rows in the group df['E'] = df.groupby('A')['D'].transform(lambda x: x[df.loc[x.index, 'B'] == 'Y'].iloc[0])
Note: The df.loc[x.index, 'B'] ensures we're only checking B values within the current group, which is safer if your data has edge cases.
Method 2: Using a Mapping Dictionary
This is a more explicit approach. First, we create a dictionary that maps each A value to its corresponding D value where B='Y', then we use this dictionary to populate column E.
# Create a map: A value -> D value where B='Y' y_d_map = df[df['B'] == 'Y'].set_index('A')['D'].to_dict() # Map each row's A value to the corresponding D value df['E'] = df['A'].map(y_d_map)
Result
Either method will give you the desired output:
| 索引 | A | B | C | D | E |
|---|---|---|---|---|---|
| 0 | 2002-01-13 15:00:00 | X | 110 | 3.9 | 1.9 |
| 1 | 2002-01-13 15:00:00 | Y | 120 | 1.9 | 1.9 |
| 2 | 2002-01-13 15:00:00 | X | 130 | 8.0 | 1.9 |
| 3 | 2002-01-13 15:00:00 | X | 140 | 9.0 | 1.9 |
| 4 | 2002-01-14 16:00:00 | X | 110 | 0.2 | 7.0 |
| 5 | 2002-01-14 16:00:00 | Y | 120 | 7.0 | 7.0 |
| 6 | 2002-01-14 16:00:00 | X | 130 | 1.6 | 7.0 |
| 7 | 2002-01-14 16:00:00 | X | 140 | 3.4 | 7.0 |
Both methods assume that each A group has exactly one row where B='Y' (which matches your sample data). If there could be multiple Y rows per group, you'd need to adjust (e.g., take the mean, first occurrence, etc.)—but based on your example, these should work perfectly.
内容的提问来源于stack exchange,提问作者Tie_24

