使用Python为分组数据创建新标识变量的实现方法
Here are a couple of straightforward ways to achieve your goal of marking the first occurrence of each name with 1 and subsequent rows with 0:
Method 1: Using duplicated()
The duplicated() method checks if a row is a duplicate of a previous row in the specified column. By negating this result and converting to integers, we get exactly the 1/0 values we need:
import pandas as pd # Your original DataFrame d = {'name': ['john', 'john', 'john', 'Tim', 'Tim', 'Tim','Bob', 'Bob'], 'Prod': ['101', '102', '101', '501', '505', '301', '302', '302'], 'Qty': ['5', '4', '1', '3', '5', '4', '1', '3']} df = pd.DataFrame(data=d) # Create the id column df['id'] = (~df['name'].duplicated()).astype(int) print(df)
Output:
name Prod Qty id 0 john 101 5 1 1 john 102 4 0 2 john 101 1 0 3 Tim 501 3 1 4 Tim 505 5 0 5 Tim 301 4 0 6 Bob 302 1 1 7 Bob 302 3 0
Method 2: Using groupby() and cumcount()
If you prefer a group-based approach, you can use groupby() on the name column and cumcount() to track the position within each group. The first row of each group will have a count of 0, so we check if the count equals 0 and convert to integer:
df['id'] = (df.groupby('name').cumcount() == 0).astype(int)
This will produce the exact same output as Method 1.
Explanation:
groupby('name')splits the DataFrame into groups for each unique name.cumcount()assigns a sequential number starting from0to each row within its group.== 0returnsTrueonly for the first row of each group.astype(int)convertsTrueto1andFalseto0.
Both methods are efficient and work well even with large datasets. The duplicated() method is slightly more concise, but the groupby() approach might be more intuitive if you're already working with grouped data.
内容的提问来源于stack exchange,提问作者singularity2047

