绘制平均占比vs占比排名(平均表现平均分)及学生学科占比分析需求
First, let's get setup with the right tools and prep our data. We'll use pandas for handling the data, matplotlib for plotting basics, and seaborn for cleaner, more insightful visuals. Assuming your input DataFrame is named df with columns for Physics, Maths, Chemistry, and observation IDs, let's start:
import pandas as pd import matplotlib.pyplot as plt import seaborn as sns # Calculate total score per student (as per your formula) df['Total'] = (df['Physics'] + df['Maths'] + df['Chemistry']) / 3 # Compute each subject's score as a percentage of the student's total score df['Physics_Pct'] = df['Physics'] / df['Total'] df['Maths_Pct'] = df['Maths'] / df['Total'] df['Chemistry_Pct'] = df['Chemistry'] / df['Total']
1. Plot: Average Subject Percentage vs. Subject Rank (By Average Performance)
This chart will show how each subject's average score占比 ranks against others, paired with their exact average percentage value.
Step-by-Step Implementation:
- Calculate the average percentage for each subject across all students
- Sort subjects by their average percentage to assign a rank (1 = highest average占比)
- Use a bar plot to make rank vs. average value comparisons straightforward
# Compute average subject percentages avg_subject_pcts = df[['Physics_Pct', 'Maths_Pct', 'Chemistry_Pct']].mean().reset_index() avg_subject_pcts.columns = ['Subject', 'Average_Percentage'] # Sort by average percentage and assign ranks avg_subject_pcts = avg_subject_pcts.sort_values(by='Average_Percentage', ascending=False) avg_subject_pcts['Rank'] = range(1, len(avg_subject_pcts)+1) # Build the plot plt.figure(figsize=(8, 5)) sns.barplot(data=avg_subject_pcts, x='Rank', y='Average_Percentage', hue='Subject', palette='viridis') # Add labels and formatting plt.title('Average Subject Score Percentage vs. Subject Rank', fontsize=14) plt.xlabel('Subject Rank (1 = Highest Average占比)', fontsize=12) plt.ylabel('Average Score Percentage (vs. Total Score)', fontsize=12) plt.ylim(0, 1) # Keep y-axis in 0-1 percentage range plt.legend(title='Subject') plt.tight_layout() plt.show()
Quick Interpretation:
You’ll immediately see which subject contributes the most to the average student’s total score. For example, if Maths holds rank 1 with an average percentage of 0.36, that means Maths makes up 36% of the typical student’s total score—more than Physics or Chemistry.
2. Plot: Subject Score Percentage vs. Normalized Total Score (0.00 = Worst, 1.00 = Best)
This visualization will reveal how subject score占比 shifts as a student’s overall performance improves or declines.
Step-by-Step Implementation:
- Normalize total scores to a 0-1 scale (so 0 = lowest total score, 1 = highest)
- Reshape the data to long format for easier plotting with seaborn
- Use a line plot with optional trendlines to highlight performance-based shifts
# Normalize total scores to 0-1 range df['Total_Normalized'] = (df['Total'] - df['Total'].min()) / (df['Total'].max() - df['Total'].min()) # Reshape data to long format (better for seaborn's hue functionality) long_df = df.melt( id_vars='Total_Normalized', value_vars=['Physics_Pct', 'Maths_Pct', 'Chemistry_Pct'], var_name='Subject', value_name='Subject_Percentage' ) # Create the plot with trendlines plt.figure(figsize=(10, 6)) sns.lineplot( data=long_df, x='Total_Normalized', y='Subject_Percentage', hue='Subject', palette='coolwarm', ci=None # Remove confidence interval if not needed; add `estimator='mean'` to show bin averages ) # Add optional regression lines to emphasize trends for subject in long_df['Subject'].unique(): subset = long_df[long_df['Subject'] == subject] sns.regplot( data=subset, x='Total_Normalized', y='Subject_Percentage', scatter=False, line_kws={'linestyle': '--'}, label=f'{subject} Trend' ) # Add labels and formatting plt.title('Subject Score Percentage vs. Overall Student Performance', fontsize=14) plt.xlabel('Normalized Total Score (0.00 = Worst, 1.00 = Best)', fontsize=12) plt.ylabel('Subject Score Percentage (vs. Student\'s Total Score)', fontsize=12) plt.ylim(0, 1) plt.legend(title='Subject & Trend') plt.tight_layout() plt.show()
Key Analysis Points:
- Trend Direction: If Maths’ trendline slopes upward, higher-performing students rely more on Maths scores for their total. If Physics’ line slopes downward, lower-performing students tend to have a larger share of their total score from Physics.
- Cross-Over Points: Watch for places where subject lines cross—this indicates a shift in which subject drives performance at different skill levels.
- Granular Binning: For deeper analysis, split the normalized total score into bins (e.g., 0-0.2, 0.2-0.4) and plot average subject percentages per bin as bar charts.
内容的提问来源于stack exchange,提问作者Alexis Drakopoulos

