如何用gnuplot绘制多类别基于Rank的累积占比曲线?
Hey there! Let's break this down into two clear, actionable steps: first calculating that cumulative proportion column using Python, then plotting the three category curves in gnuplot.
Step 1: Calculate Cumulative Proportions with Python
Since you're comfortable with bash but need to handle this data calculation, Python's pandas library makes this process quick and clean. Here's a script that reads your raw data, computes the cumulative percentage for each category, and saves the result to a new file:
import pandas as pd # Load your input data (update the filename to match your actual file) # Assuming your input has no header row; adjust names if your data has headers df = pd.read_csv("raw_data.csv", header=None, names=["Rank", "Category"]) # First, get the total number of entries for each category total_per_category = df["Category"].value_counts() # Calculate cumulative count per category (add 1 because cumcount starts at 0) df["Cumulative_Pct"] = df.groupby("Category").cumcount() + 1 # Divide by total to get the proportion (0-1 range) df["Cumulative_Pct"] = df["Cumulative_Pct"] / df["Category"].map(total_per_category) # Save the result to a new CSV (this will have all three columns) df.to_csv("processed_data.csv", index=False)
Quick breakdown of the logic:
groupby("Category").cumcount()tracks how many times the current category has appeared before the current row; adding 1 gives the total count up to and including the current row.map(total_per_category)looks up the total number of entries for each category, so dividing the cumulative count by this total gives the cumulative proportion you need.
Step 2: Plot the Curves with Gnuplot
Once you have the processed CSV with all three columns, you can use gnuplot to plot each category as a distinct curve. Here's a sample gnuplot script tailored to your needs:
# Set output to a PNG file (adjust size/format as needed) set terminal pngcairo enhanced font 'Arial,12' size 1000,700 set output 'category_cumulative_plot.png' # Configure axis labels and plot title set xlabel "Rank" font 'Arial,14' set ylabel "Cumulative Proportion" font 'Arial,14' set title "Cumulative Distribution of Categories by Rank" font 'Arial,16' # Set y-axis to range from 0 to 1 (matches your proportion values) set yrange [0:1] set grid ytics lt 0 lw 1 lc rgb '#eeeeee' # Plot each category as a distinct line plot 'processed_data.csv' using 1:($2 eq "type1" ? $3 : NaN) \ title 'Type 1' with lines lt 1 lc rgb '#1f77b4' lw 2, \ '' using 1:($2 eq "type2" ? $3 : NaN) \ title 'Type 2' with lines lt 1 lc rgb '#ff7f0e' lw 2, \ '' using 1:($2 eq "type3" ? $3 : NaN) \ title 'Type 3' with lines lt 1 lc rgb '#2ca02c' lw 2
Key details in the plot script:
- The
using 1:($2 eq "type1" ? $3 : NaN)clause filters rows to only include those where the category is "type1", using Rank (column 1) as the X-axis and Cumulative_Pct (column 3) as the Y-axis. Rows that don't match are ignored (marked as NaN). ltsets line type,lc rgbsets a unique color for each category, andlwincreases line width for better readability.- The subtle grid lines make it easier to interpret the y-axis proportion values.
To run this, save the script as plot_script.gp and execute it with gnuplot plot_script.gp in your terminal.
内容的提问来源于stack exchange,提问作者B. du Garay

