使用Pandas绘制分类数据汇总散点图的方法求助
Got it, let's tackle this scatter plot problem step by step. Here's how you can achieve exactly what you're looking for:
Step 1: Reshape Your Data to Long Format
Your current DataFrame is in "wide" format—each row has multiple observations across columns. To plot each data point individually, we need to convert it to "long" format using pd.melt(). This will create a row for every combination of feature (wild/raw/rotten) and its corresponding level.
Step 2: Define Ordered Categories for the Y-Axis
By default, string values are sorted alphabetically, which would put "little" at the top of your y-axis. We need to explicitly set the order of your levels ("very" > "medium" > "little") using pandas' categorical data type.
Step 3: Plot the Scatter Plot
Using seaborn (or matplotlib) will make it easy to handle categorical axes with our ordered levels.
Here's the full code implementation:
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt # Your original DataFrame df = pd.DataFrame({ 'wild': ['little', 'little', 'very'], 'raw': ['medium', 'medium', 'very'], 'rotten': ['little', 'very', 'medium'] }) # Convert to long format melted_df = df.melt(var_name='feature', value_name='level') # Set ordered categories for the level column melted_df['level'] = pd.Categorical( melted_df['level'], categories=['very', 'medium', 'little'], ordered=True ) # Create the scatter plot plt.figure(figsize=(8, 5)) sns.scatterplot(data=melted_df, x='feature', y='level', s=120, color='#1f77b4') # Add labels and title plt.title('Distribution of Feature Levels', fontsize=14) plt.xlabel('Feature', fontsize=12) plt.ylabel('Level', fontsize=12) plt.show()
What This Does:
- The
melt()function transforms your 3-row DataFrame into a 9-row DataFrame, where each row represents one data point (e.g., "wild" with value "little"). - The ordered categorical ensures the y-axis displays "very" at the top, followed by "medium" and "little"—exactly the order you wanted.
- The scatter plot will show all 9 points, grouped by the feature on the x-axis and aligned to the correct level on the y-axis.
If you prefer using matplotlib directly instead of seaborn, you can map the categories to numerical values and set custom y-ticks:
# Alternative matplotlib-only approach level_mapping = {'very': 2, 'medium': 1, 'little': 0} melted_df['y_value'] = melted_df['level'].map(level_mapping) plt.figure(figsize=(8,5)) plt.scatter(melted_df['feature'], melted_df['y_value'], s=120) # Set custom y-ticks and labels plt.yticks([2,1,0], ['very', 'medium', 'little']) plt.title('Distribution of Feature Levels') plt.xlabel('Feature') plt.ylabel('Level') plt.show()
Both approaches will give you the scatter plot you're aiming for.
内容的提问来源于stack exchange,提问作者pitosalas

