使用Python计算分组数据的极差与四分位数并通过Matplotlib绘制可视化图表
Got it, let's tackle this problem step by step. Since we're working with grouped frequency data (no individual values), we'll use standard interpolation methods for quartiles and lean on midpoints for visualization. Here's a complete Python implementation:
Step 1: Set Up the Data
First, we'll define our intervals, frequencies, and calculate key values like midpoints and cumulative frequencies (needed for quartile calculations).
import numpy as np import matplotlib.pyplot as plt # Define our grouped score data score_intervals = ["0-10", "10-20", "20-30", "30-40", "40-50"] frequencies = [5, 13, 20, 32, 60] # Calculate midpoints of each interval (for visualization and weighted calculations) midpoints = [(int(interval.split("-")[0]) + int(interval.split("-")[1])) / 2 for interval in score_intervals] # Compute cumulative frequencies to find quartile positions cumulative_freqs = np.cumsum(frequencies) total_observations = cumulative_freqs[-1] # Total number of data points
Step 2: Calculate the Range
For grouped data, the standard range is the upper bound of the highest interval minus the lower bound of the lowest interval. Alternatively, you could use midpoint range, but the first method is more widely accepted for grouped datasets.
# Extract bounds from intervals lowest_bound = int(score_intervals[0].split("-")[0]) highest_bound = int(score_intervals[-1].split("-")[1]) data_range = highest_bound - lowest_bound print(f"Range of the score data: {data_range}") # Output: Range of the score data: 50
Step 3: Calculate Quartiles (Q1, Q2, Q3)
Since we don't have individual values, we use interpolation to find quartiles. The formula for the k-th quartile is:
Qk = L + ((k*N/4 - F)/f) * w
Where:
- L = Lower bound of the interval containing the quartile
- N = Total number of observations
- F = Cumulative frequency of the interval before the quartile's interval
- f = Frequency of the quartile's interval
- w = Width of the interval (all intervals here are 10, so w=10)
Let's code this into a reusable function:
def get_grouped_quartile(k, cumulative_freqs, freqs, intervals): """Calculate the k-th quartile (k=1,2,3) for grouped frequency data""" quartile_position = k * total_observations / 4 # Find which interval contains the quartile for idx, cum_freq in enumerate(cumulative_freqs): if cum_freq >= quartile_position: # Extract interval details lower_bound = int(intervals[idx].split("-")[0]) prev_cum_freq = cumulative_freqs[idx-1] if idx > 0 else 0 interval_freq = freqs[idx] interval_width = int(intervals[idx].split("-")[1]) - lower_bound break # Apply interpolation formula quartile = lower_bound + ((quartile_position - prev_cum_freq) / interval_freq) * interval_width return quartile # Calculate Q1 (25th percentile), Q2 (median), Q3 (75th percentile) q1 = get_grouped_quartile(1, cumulative_freqs, frequencies, score_intervals) q2 = get_grouped_quartile(2, cumulative_freqs, frequencies, score_intervals) q3 = get_grouped_quartile(3, cumulative_freqs, frequencies, score_intervals) print(f"Q1: {round(q1, 2)}") print(f"Median (Q2): {round(q2, 2)}") print(f"Q3: {round(q3, 2)}") # Output: # Q1: 23.12 # Median (Q2): 35.94 # Q3: 43.33
Step 4: Visualize the Data
We'll use two plots to summarize the data:
- A histogram to show the frequency distribution across intervals.
- A box plot to visualize the range, quartiles, and median (using weighted midpoints to simulate individual data points).
Histogram for Frequency Distribution
# Convert intervals to numerical bin edges for the histogram bin_edges = [int(interval.split("-")[0]) for interval in score_intervals] bin_edges.append(int(score_intervals[-1].split("-")[1])) # Add final upper bound plt.figure(figsize=(10, 6)) plt.hist([midpoints]*frequencies, bins=bin_edges, edgecolor="black") plt.title("Frequency Distribution of Scores") plt.xlabel("Score Intervals") plt.ylabel("Number of Students") plt.xticks(bin_edges) plt.grid(axis="y", linestyle="--", alpha=0.7) plt.show()
Box Plot for Quartile & Range Summary
# Create a weighted list of midpoints to simulate individual data points weighted_scores = [] for mid, freq in zip(midpoints, frequencies): weighted_scores.extend([mid]*freq) plt.figure(figsize=(8, 6)) box_plot = plt.boxplot(weighted_scores, vert=False, patch_artist=True) plt.title("Box Plot of Score Distribution") plt.xlabel("Score") # Customize the plot for readability for patch in box_plot["boxes"]: patch.set_facecolor("#88c999") plt.grid(axis="x", linestyle="--", alpha=0.7) # Annotate key statistics plt.annotate(f"Q1: {round(q1,2)}", (q1, 1.1)) plt.annotate(f"Median: {round(q2,2)}", (q2, 1.1)) plt.annotate(f"Q3: {round(q3,2)}", (q3, 1.1)) plt.annotate(f"Range: {data_range}", (lowest_bound, 0.9)) plt.annotate(f"Max Score: {highest_bound}", (highest_bound, 0.9)) plt.show()
Quick Notes
- The histogram highlights that the largest group of scores falls in the 40-50 range.
- The box plot gives a compact overview of the data's spread: it shows the minimum (0), Q1, median, Q3, and maximum (50) values at a glance.
内容的提问来源于stack exchange,提问作者om9595

