按列值汇总Pandas DataFrame:二进制指标行分组统计需求
To solve this problem—grouping identical binary rows and adding a count column—we can leverage Pandas' built-in value_counts() method, which efficiently counts occurrences of unique rows. Here's a step-by-step breakdown:
Step 1: Understand the Goal
We need to take a DataFrame with binary columns, identify all unique row combinations, count how many times each unique row appears, and create a new DataFrame that pairs these unique rows with their respective counts.
Step 2: Use value_counts() with sort=False
The value_counts() method directly counts unique row combinations out of the box. Setting sort=False ensures we preserve the order of the first occurrence of each unique row (matching the order in your example output).
Step 3: Convert to a DataFrame and Rename the Count Column
After generating the counts, we use reset_index() to convert the resulting Series back into a structured DataFrame, then rename the count column for clarity.
Full Code Example
import pandas as pd # Your input DataFrame df = pd.DataFrame([ [0,1,1,0], [0,1,1,0], [0,0,0,1], [0,0,0,1], [1,1,1,0], [1,1,1,1], [1,1,1,0] ]) # Generate the result DataFrame res = df.value_counts(sort=False).reset_index(name='count') # Optional: If you prefer the count column to be unnamed (Pandas recommends named columns for clarity) # res.rename(columns={'count': ''}, inplace=True) print(res)
Output
This produces exactly the result you specified:
0 1 2 3 count 0 0 1 1 0 2 1 0 0 0 1 2 2 1 1 1 0 2 3 1 1 1 1 1
Alternative Approach Using groupby
If you prefer using groupby instead, here's an equivalent solution (note: groupby sorts groups by default, so to match the example's order you'd need extra steps to track first occurrence):
res = df.groupby(df.columns.tolist()).size().reset_index(name='count')
The value_counts() method is more concise and efficient for this use case, especially when preserving the order of first occurrence is important.
内容的提问来源于stack exchange,提问作者Research100

