Groupby未识别全部ID:Pandas分组求和结果不符合预期
Hey Peter, let's break down what's going on here—your ID cleanup actually worked perfectly (both 11 and 45 are showing up in the grouped output), so the issue isn't with ID recognition. The problem lies in how you're aggregating your data.
Looking at your raw data:
campaignid name isb
11 aaa true
11 asa false
11 afs true
45 adfs false
45 ads true
When you run groupby(["campaignid"]).sum(), it adds up the isb values. Since true maps to 1 and false maps to 0, campaign 11 gives 1 + 0 + 1 = 2—this matches your current output. But your expected result of 3 means you want to count all rows per campaign, not just the number of true entries.
Solution: Count total rows per campaign
Swap out sum() for size() to get the total number of records in each group, then reset the index to match your desired format:
# Keep your ID cleanup step—it's correct cw["campaignid"] = cw["campaignid"].str.replace(r'\ ', '') # Group by campaignid and count rows, rename the result column to 'leads' grouped_cw = cw.groupby(["campaignid"]).size().reset_index(name='leads')
Bonus: If you actually wanted to count true values
If your expectation was a typo and you do need the count of true entries, first make sure your isb column is properly typed (if it's stored as strings instead of booleans):
cw["campaignid"] = cw["campaignid"].str.replace(r'\ ', '') # Convert string 'true'/'false' to boolean values cw['isb'] = cw['isb'].map({'true': True, 'false': False}) # Sum the boolean values (True=1, False=0) grouped_cw = cw.groupby(["campaignid"])['isb'].sum().reset_index(name='leads')
Either way, your ID space-removal step was working as intended—you just needed to adjust the aggregation method to match your desired output.
内容的提问来源于stack exchange,提问作者Peter

