基于Pandas DataFrame处理分类数据:两项可视化与统计任务问询
Hey there! Let's work through your two Pandas DataFrame tasks together—they're straightforward once you break them down. I'll use Pandas for data manipulation and Matplotlib for plotting (since it's the most common pairing for this kind of visualization).
First, we'll group the data by the degree of endangerment column, calculate the average number of speakers per group, then plot it with a logarithmic y-axis to handle any large differences in speaker counts.
Here's the code with comments to explain each step:
import pandas as pd import matplotlib.pyplot as plt # Assuming your DataFrame is named 'df' # 1. Group by endangerment degree and calculate mean speaker count grouped_mean = df.groupby('degree of endangerment')['Number of speakers'].mean().sort_values() # 2. Create the bar chart plt.figure(figsize=(10, 6)) bars = plt.bar(grouped_mean.index, grouped_mean.values) # 3. Set log scale for y-axis (per your request) plt.yscale('log') # 4. Add labels and title for clarity plt.xlabel('Degree of Endangerment') plt.ylabel('Mean Number of Speakers (Log Scale)') plt.title('Mean Speaker Count by Language Endangerment Degree') # Optional: Add value labels on top of each bar for readability for bar in bars: height = bar.get_height() plt.text(bar.get_x() + bar.get_width()/2., height, f'{height:.2f}', ha='center', va='bottom') plt.xticks(rotation=45, ha='right') # Rotate x-labels to prevent overlap plt.tight_layout() # Adjust layout to fit all elements plt.show()
Pro Tip: If your dataset has missing values in either column, add df = df.dropna(subset=['degree of endangerment', 'Number of speakers']) before grouping to avoid skewed results.
To get full descriptive stats (like mean, median, standard deviation, min/max, etc.) for each group, we can use Pandas' built-in describe() method after grouping.
Full Descriptive Stats
# Group by endangerment degree and generate comprehensive descriptive statistics descriptive_stats = df.groupby('degree of endangerment')['Number of speakers'].describe() # Print or display the results print(descriptive_stats) # If you're in a Jupyter Notebook, just type `descriptive_stats` to see a formatted table
Custom Stats (If You Only Need Specific Metrics)
If you don't need all the stats from describe(), you can specify exactly which ones you want using agg():
custom_stats = df.groupby('degree of endangerment')['Number of speakers'].agg( mean='mean', median='median', standard_deviation='std', min_speakers='min', max_speakers='max', total_languages='count' ) print(custom_stats)
This will give you a clean, structured table of the metrics that matter most to you.
内容的提问来源于stack exchange,提问作者Mark

