DataFrame按数值区间分组统计并绘图的问题排查与实现需求
Hey there! Let's break down why your current code is returning that unexpected 22 count, then walk through the correct way to group your data and create the visualization you need.
What's Wrong With the Original Code?
You hit two key issues here:
- Boolean Operator Precedence: The
&operator has higher priority than comparison operators (>,<=). So your conditiondf['count'] > 0 & (df['count'] <= 2)gets parsed asdf['count'] > (0 & (df['count'] <= 2))—which simplifies to checking if values are greater than 0 (all your entries are!). That's why you're getting 22 (the total number of rows in your DataFrame). - Incorrect Counting Method: Using
count()on a filtered DataFrame returns the number of non-null values per column. Since your filtered DataFrame still has all 22 rows (thanks to the precedence bug), it's counting all entries instead of the number of rows matching your criteria.
Correct Approach: Grouping & Counting
The cleanest way to handle this grouping is with pd.cut—it's designed exactly for binning numerical values into intervals. Here's how to do it step by step:
Step 1: Define Bins & Labels
First, set up your intervals and corresponding group labels:
import pandas as pd # Original data data = {'count':[11, 113, 53, 416, 3835, 143, 1, 1, 1, 2, 3, 4, 3, 4, 4, 6, 7, 7, 8,8,8,9]} df = pd.DataFrame(data) # Define bins: (0,2] = 1-2, (2,4] =3-4, etc. bins = [0, 2, 4, 6, 8, 10, float('inf')] group_labels = ['1-2', '3-4', '5-6', '7-8', '9-10', '>10']
Step 2: Assign Groups & Count Entries
Add a grouping column to your DataFrame, then count how many entries fall into each group:
# Assign each row to a group df['group'] = pd.cut(df['count'], bins=bins, labels=group_labels, right=True) # Count entries per group, preserving the order of your labels group_counts = df['group'].value_counts().reindex(group_labels) print(group_counts)
This will output the correct counts:
1-2 4 3-4 5 5-6 1 7-8 5 9-10 1 >10 6 Name: group, dtype: int64
If You Prefer Manual Filtering (Without pd.cut)
If you want to stick to manual boolean filtering, fix the precedence issue and use len() to count rows:
# Correct way to count 1-2 group count_1_2 = len(df[(df['count'] >= 1) & (df['count'] <= 2)]) print(count_1_2) # Outputs 4, which is correct (three 1s and one 2)
Visualizing the Grouped Data
Now that you have the correct counts, you can easily plot them using matplotlib or seaborn. Here's an example with seaborn for a clean look:
import matplotlib.pyplot as plt import seaborn as sns plt.figure(figsize=(10, 6)) sns.barplot(x=group_counts.index, y=group_counts.values, palette='viridis') plt.title('Distribution of Count Values by Group') plt.xlabel('Value Intervals') plt.ylabel('Number of Entries') plt.grid(axis='y', alpha=0.3) plt.show()
Or use pandas' built-in plotting for simplicity:
group_counts.plot(kind='bar', figsize=(10,6), color='#429bf5') plt.title('Distribution of Count Values by Group') plt.xlabel('Value Intervals') plt.ylabel('Number of Entries') plt.show()
内容的提问来源于stack exchange,提问作者Exa

