如何基于DataFrame变量组合绘制整合式散点图?
Hey Megan, I've got you covered! Let's optimize your code so you can visualize all your required scatter plot combinations automatically—no manual parameter tweaking needed. Here's a practical, clean solution:
Refactored Code for Automatic Multi-Combination Plotting
We'll use a dictionary to map your kind labels to their DataFrames, generate all valid index pairs automatically, and use subplots to display all combinations in one organized figure.
import pandas as pd import matplotlib.pyplot as plt import numpy as np from itertools import combinations # --- Test DataFrames --- food = {'col1': [1, 2], 'col2': [3, 4]} food_df = pd.DataFrame(data=food) meat = {'col1': [3, 6], 'col2': [1, 2]} meat_df = pd.DataFrame(data=meat) vege = {'col1': [5, 9], 'col2': [0, 3]} vege_df = pd.DataFrame(data=vege) # --- Core Parameters --- kind = ["food", "meat", "vege"] features = ["Attack", "Volume"] col = ["dd", "aa"] # Map kind labels to their DataFrames (avoids messy hardcoding later) df_mapping = {label: df for label, df in zip(kind, [food_df, meat_df, vege_df])} # Generate all valid kind index pairs automatically (no manual (0,1) entries!) kind_index_pairs = list(combinations(range(len(kind)), 2)) # Assign unique colors to each feature for clear differentiation feature_colors = {"Attack": "#1f77b4", "Volume": "#ff7f0e"} # --- Create Subplot Layout --- # 1 row, 3 columns to fit all 3 variable combinations fig, axes = plt.subplots(1, 3, figsize=(18, 6)) # Iterate over each kind pair and corresponding subplot for (idx1, idx2), ax in zip(kind_index_pairs, axes.flat): # Grab the two DataFrames for this combination df_x = df_mapping[kind[idx1]] df_y = df_mapping[kind[idx2]] # Plot data for each feature for feature in features: # Get the row index for the current feature feat_row_idx = features.index(feature) # Extract the data points for this feature x_data = df_x.iloc[feat_row_idx].values y_data = df_y.iloc[feat_row_idx].values # Plot scatter points with unique color ax.scatter(x_data, y_data, color=feature_colors[feature], label=feature, s=100, alpha=0.8) # Add annotations for each column label ("dd", "aa") for point_idx, txt in enumerate(col): ax.annotate(txt, (x_data[point_idx], y_data[point_idx]), xytext=(x_data[point_idx]+0.1, y_data[point_idx]+0.1), fontsize=10, weight="bold") # Configure subplot labels and title ax.set_xlabel(kind[idx1], fontsize=12) ax.set_ylabel(kind[idx2], fontsize=12) ax.set_title(f"{kind[idx1]} vs {kind[idx2]}", fontsize=14, pad=15) ax.legend(title="Feature", fontsize=10) ax.grid(True, alpha=0.3) # Adjust spacing between subplots for better readability plt.tight_layout() plt.show()
Key Improvements Breakdown:
- Flexible Data Mapping: The
df_mappingdictionary lets you quickly access any DataFrame using itskindlabel—adding new DataFrames later only requires updating this dictionary and thekindlist. - Automatic Pair Generation:
itertools.combinationscreates all valid (kind1, kind2) index pairs automatically, so you never have to manually adjustkind1/kind2values again. - Organized Subplot Layout: All three scatter plot combinations are displayed in one figure, making it easy to compare relationships between variables at a glance.
- Clear Visuals: Unique colors for each feature and bold annotations ensure your data points are easy to identify and interpret.
- Maintainable Code: The modular structure means you can tweak colors, point sizes, or add new features/DataFrames without rewriting large sections of code.
If you'd prefer to plot all combinations on a single axis instead of subplots, we could adjust the code to use different markers for each kind pair—but subplots are generally the clearer choice for comparing multiple variable relationships.
内容的提问来源于stack exchange,提问作者Megan
相关产品推荐
相关产品推荐

