Pandas按A列分组求和B列并保存为新列B_updated的实现问题
Got it, the issue you're facing is that groupby('A')['B'].sum() gives you a condensed Series with one row per unique 'A' value, which isn't aligned with your original DataFrame for further operations. The fix is to use transform() instead—it calculates the grouped sum and broadcasts the result back to every row in the original DataFrame, perfect for creating your new "B_updated" column.
Here's how to do it with your data:
First, let's recreate your original DataFrame for clarity:
import pandas as pd data = {'A': [61880, 62646, 62651, 62656, 62783, 61880, 62646], 'B': [7, 8, 9, 10, 11, 3, 2]} df1 = pd.DataFrame(data)
Now apply the transform method to add the new column:
df1['B_updated'] = df1.groupby('A')['B'].transform('sum')
This will give you a DataFrame where each row has the sum of 'B' values for its corresponding 'A' in the new column:
A B B_updated 0 61880 7 10 1 62646 8 10 2 62651 9 9 3 62656 10 10 4 62783 11 11 5 61880 3 10 6 62646 2 10
Why this works:
groupby('A')['B']groups your data by the 'A' column.transform('sum')computes the sum for each group, then maps that sum back to every row in the original group. Unlike.sum()which returns a single row per group, transform preserves the original row count so you can easily assign it as a new column.
Now you can perform any column-wise operations you need using "B_updated" alongside your original columns!
内容的提问来源于stack exchange,提问作者Nazar Tarlanli

