如何按索引对DataFrame多列绘制值频率散点图?求技术协助
Let me break down what's going wrong with your current code and fix it to get the frequency scatter plot you want.
First, your DataFrame columns are stored as strings instead of numeric values — that's a key issue. Second, your current plt.plot call is just plotting each row's raw value against the index (with connecting lines), which isn't a frequency plot at all. A frequency scatter plot should visualize how often each unique value appears in a column.
Here are two common interpretations of what you might be aiming for, with working code for both:
Option 1: Scatter plot of value vs. its frequency (per column)
This shows each unique value in a column on the x-axis, and how many times it appears (frequency) on the y-axis:
import pandas as pd import matplotlib.pyplot as plt # Your original dataset df = pd.DataFrame({'A': ['0','0','0','2','0','0','0','3'], 'B': ['0','0','0','4','3','0','0','4'], 'C': ['1','1','2','8','9','6','1','10']}) # Step 1: Convert columns from strings to integers (critical for proper sorting/plotting) df = df.astype(int) # Step 2: Calculate frequency counts for each column, sorted by value freq_A = df['A'].value_counts().sort_index() freq_B = df['B'].value_counts().sort_index() freq_C = df['C'].value_counts().sort_index() # Step 3: Plot the frequency scatter points plt.figure(figsize=(10, 5), dpi=80) # Column A: red x markers plt.scatter(freq_A.index, freq_A.values, color='red', marker='x', label='Column A') # Column B: green triangle markers plt.scatter(freq_B.index, freq_B.values, color='green', marker='^', label='Column B') # Column C: blue square markers plt.scatter(freq_C.index, freq_C.values, color='blue', marker='s', label='Column C') # Add labels and legend for clarity plt.xlabel('Value') plt.ylabel('Frequency (Number of Occurrences)') plt.title('Frequency of Unique Values by Column') plt.legend() plt.show()
Option 2: Scatter plot of index vs. frequency of the value at that index
If you want to plot, for each row index, how often the value in that row's column appears across the entire column (e.g., row 3 in column A has value 2, which appears once — so y=1 at x=3), use this code:
import pandas as pd import matplotlib.pyplot as plt df = pd.DataFrame({'A': ['0','0','0','2','0','0','0','3'], 'B': ['0','0','0','4','3','0','0','4'], 'C': ['1','1','2','8','9','6','1','10']}) df = df.astype(int) # Create dictionaries mapping each value to its frequency in the column freq_map_A = df['A'].value_counts().to_dict() freq_map_B = df['B'].value_counts().to_dict() freq_map_C = df['C'].value_counts().to_dict() # Add new columns with the frequency of each row's value df['A_freq'] = df['A'].map(freq_map_A) df['B_freq'] = df['B'].map(freq_map_B) df['C_freq'] = df['C'].map(freq_map_C) # Plot frequency against index plt.figure(figsize=(10, 5), dpi=80) plt.scatter(df.index, df['A_freq'], color='red', marker='x', label='Column A') plt.scatter(df.index, df['B_freq'], color='green', marker='^', label='Column B') plt.scatter(df.index, df['C_freq'], color='blue', marker='s', label='Column C') plt.xlabel('Row Index') plt.ylabel('Frequency of the Value at This Index') plt.title('Frequency of Column Values by Row Index') plt.legend() plt.show()
Both options fix the core issues with your original code: converting string values to numeric, and focusing on frequency rather than just plotting raw values. Pick the one that matches your actual goal!
内容的提问来源于stack exchange,提问作者naura tsalits

