Pandas按Type和SubType分组后筛选指定Type并统计Municipality频次的问题
Hey there! I see exactly what's tripping you up here—your current code is grouping by the boolean result of Type == 'Social Media' instead of first narrowing down to only those rows. That's why you're getting the extra True/False sections in your output. Let's fix this together.
Why Your Current Code Isn't Working
When you pass df['Type'] == 'Social Media' directly into groupby(), Pandas creates two separate groups: one for rows where the condition is true (your desired Social Media entries) and one where it's false (all other rows). That's the reason you're seeing the unwanted False section with non-Social Media data.
The Simple Fix: Filter First, Then Group
The most straightforward way to get your desired result is to filter the dataframe to only include Social Media rows first, then do your grouping and counting. Here's the code:
# Step 1: Filter to keep only Social Media entries social_media_only = df[df['Type'] == 'Social Media'] # Step 2: Group by Type + SubType, then count Municipality occurrences result = social_media_only.groupby(['Type', 'SubType'])['Municipality'].value_counts() # Optional: Reset index to turn it into a flat dataframe with a "Count" column result = result.reset_index(name='Count')
Alternative: Group First, Then Filter
If you prefer to group the entire dataframe first and then narrow it down, you can use .query() to keep only the Social Media groups:
# Group all data by Type + SubType, then count Municipality values grouped_data = df.groupby(['Type', 'SubType'])['Municipality'].value_counts() # Filter to retain only Social Media groups result = grouped_data.query("Type == 'Social Media'")
What Your Result Will Look Like
Either method will give you exactly the output you're aiming for. If you reset the index, it'll look like this:
| Type | SubType | Municipality | Count |
|---|---|---|---|
| Social Media | New Castle | 2 | |
| Social Media | San Andreas | 1 | |
| Social Media | New Castle | 1 | |
| Social Media | Tiktok | San Andreas | 1 |
Or in the multi-index format (without resetting the index):
Type SubType Municipality Social Media Facebook New Castle 2 San Andreas 1 Instagram New Castle 1 Tiktok San Andreas 1 Name: Municipality, dtype: int64
This will get rid of the unwanted False group and leave you with only the Social Media statistics you need.
备注:内容来源于stack exchange,提问作者Chanel Loski

